VLDB 2026 Research / reviewers in the wild / expert
Jianrong Zhang
dblp:119/5950
· DBLP profile ↗
16ranked-venue papers
6as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 5 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 6 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Live Demonstration: An Agile FPGA-Overlayed CGRA SoC for High-Efficiency Computing
Jiahang Lou, Jianrong Zhang, Yuan Dai, Zewei Zhong, Wenbo Yin, Lingli Wang |
ISCAS | 2 |
| 2026 | Adaptive Bitrate Live Streaming over HTTP-FLV: A Practical System Perspective
Tong Meng, Bingcong Lu, Jinghao Yuan, Huanting Liu, Nailiang Wu, Zhou Sha, Changqing Yan, Jianrong Zhang, Jianxin Kuang, Li Song 0001 |
SIGCOMM | 11 |
| 2026 | GRAG-ProSafe QAS: Graph retrieval-augmented generation for process production safety intelligent question and answer system
Jianrong Zhang, Qingwen Wei, Huayu Zhong |
Expert Syst. Appl. | 1 |
| 2026 | BeatDance: Generating beat-consistent 3D dance with hierarchical spatial-temporal modeling
Xiaojian Shen, Dahu Shi, Jianrong Zhang, Yunzhi Zhuge, Zhiliang Wu, Guanghui Yue 0001, Wei Zhou 0021 |
Pattern Recognit. | 3 |
| 2025 | Heterogeneous Graph Neural Networks with Ordinal Regression for Legal Case Retrieval
Jianrong Zhang, Xiaoyu Kang, Zhixin Shi |
IEEE Big Data | 1 |
| 2025 | EnergyMoGen: Compositional Human Motion Generation with Energy-Based Diffusion Model in Latent SpaceabstractDiffusion models, particularly latent diffusion models, have demonstrated remarkable success in text-driven human motion generation. However, it remains challenging for latent diffusion models to effectively compose multiple semantic concepts into a single, coherent motion sequence. To address this issue, we propose EnergyMoGen, which includes two spectrums of Energy-Based Models: ❶ We interpret the diffusion model as a latent-aware energy-based model that generates motions by composing a set of diffusion models in latent space; ❷ We introduce a semantic-aware energy model based on cross-attention, which enables semantic composition and adaptive gradient descent for text embeddings. To overcome the challenges of semantic inconsistency and motion distortion across these two spectrums, we introduce Synergistic Energy Fusion. This design allows the motion latent diffusion model to synthesize high-quality, complex motions by combining multiple energy terms corresponding to textual descriptions. Experiments show that our approach outperforms existing state-of-the-art models on various motion generation tasks, including text-to-motion generation, compositional motion generation, and multi-concept motion generation. Additionally, we demonstrate that our method can be used to extend motion datasets and improve the text-to-motion task. Jianrong Zhang, Hehe Fan, Yi Yang 0001 |
CVPR | 1 |
| 2025 | LEMOE: LLM-Enhanced Multi-Objective Bayesian Optimization for Microarchitecture ExplorationabstractDesigning processor microarchitectures is increasingly challenging due to a vast design space and the need to balance multiple metrics. Traditional algorithm-driven design space exploration (DSE) approaches often struggle to incorporate the extensive domain knowledge of expert architects. To address this, we introduce LEMOE, a multi-objective microarchitecture optimization framework that leverages large language model (LLM) to enhance an implicit Bayesian model. LEMOE features a program-aware warm-up phase utilizing LLM and LLVM to produce an initial design set with rich prior knowledge. By harnessing LLM’s contextual learning, our approach improves surrogate modeling and sampling under sparse data conditions. Experiment results show that LEMOE achieves a $22.8 \%$ improvement in energy efficiency with the same number of iterations and a $2.9 \times$ runtime speedup for the same target compared to prior works. Jingyuan Li 0003, Jianrong Zhang, Wenbo Yin, Lingli Wang |
DAC | 2 |
| 2025 | Detecting and Adapting to Stealthy Label-Inversion Drifts via Conditional Distribution InferenceabstractDeep learning (DL) based malicious traffic detectors have been widely developed to detect diverse network attacks, yet they are suffering from significant performance degradation due to concept drift. Existing anti-concept drift arts focus on combating the drifting traffic whose features significantly diverge from training traffic. However, they neglect a stealthy yet common situation where the testing traffic has similar features to the training traffic but with opposite ground truth labels. As a result, the DL-based detectors would always make incorrect predictions for the stealthy drifting traffic, insufficient to perform long-term real-world intrusion detection. In this paper, we propose Chameleon, a novel active learning framework that combats stealthy drifting traffic by inferring the conditional distribution of the testing traffic with small manual labeling overhead. Specifically, Chameleon measures the fine-grained correlations between the high-dimensional and heterogeneous testing traffic and selects a small number of highly representative testing traffic samples for manual labeling, to accurately infer other testing samples’ labels. With the inferred labels, Chameleon checks the conditional distribution shift from the training to testing traffic to detect concept drift and incrementally trains the DL-based detectors to make them effectively adapt to the shifted distribution. Extensive experiments with six supervised and unsupervised DL-based detectors on three public and four synthetic datasets show that, under stealthy drifting traffic, Chameleon improves the AUT of the DL-based detectors by a range of $18.53 \%$ to $23.89 \%$, while the improvement of SOTA baselines is only between $0.06 \%$ and $1.86 \%$. Xiaoli Zhang 0003, Qilei Yin, Jianrong Zhang, Ke Xu 0002, Qi Li 0002, Xu-Cheng Yin |
RAID | 6 |
| 2025 | Who are querying for me? Measuring the dependency and centralization in recursive resolution
Qiuyun Wang, Jianrong Zhang, Baojiang Cui, Zhengwei Jiang |
Comput. Secur. | 3 |
| 2025 | Protein Captioning: Bridging the Gap between Protein Sequences and Natural LanguagesabstractWe introduce the multimodal task of Protein Captioning , which is an easy-to-understand and flexible way for protein analysis. Compared to specific protein recognition or classification tasks, such as enzyme reaction classification and gene ontology term prediction, protein captioning provides comprehensive textural descriptions for proteins, thus playing a key role in bridging the gap between protein sequences and natural languages. To address the problem, we propose a simple yet effective method, Protein-to-Text Generative Pre-Trained Transformer (P2T-GPT), to fuse multimodal embeddings and translate the chain of amino acid residues in a protein to a sequence of natural language words, i.e., text. For the evaluation of protein captioning, we collect the ProteinCap dataset that contains 94,454 protein-text pairs. Experiments on ProteinCap demonstrate the effectiveness of the proposed P2T-GPT on protein captioning. For example, our method obtains improvements of 8.74, 10.03, and 11.05 in the BERTScore compared to the baseline model on ProteinCap- \(\alpha,\beta,\gamma\) , respectively. As minor contributions, first, P2T-GPT provides a way to connect protein science and Large Language Models (LLMs). By appending ChatGPT, our method can interact in a conversational way to answer questions given a protein. Second, we show that protein captioning can be treated as a pre-trained task that can benefit a range of downstream tasks, to a certain extent. Jianrong Zhang, Hehe Fan, Yi Yang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | Hand-Centric Motion Refinement for 3D Hand-Object Interaction via Hierarchical Spatial-Temporal ModelingabstractHands are the main medium when people interact with the world. Generating proper 3D motion for hand-object interaction is vital for applications such as virtual reality and robotics. Although grasp tracking or object manipulation synthesis can produce coarse hand motion, this kind of motion is inevitably noisy and full of jitter. To address this problem, we propose a data-driven method for coarse motion refinement. First, we design a hand-centric representation to describe the dynamic spatial-temporal relation between hands and objects. Compared to the object-centric representation, our hand-centric representation is straightforward and does not require an ambiguous projection process that converts object-based prediction into hand motion. Second, to capture the dynamic clues of hand-object interaction, we propose a new architecture that models the spatial and temporal structure in a hierarchical manner. Extensive experiments demonstrate that our method outperforms previous methods by a noticeable margin. Yuze Hao, Jianrong Zhang, Tao Zhuo, Fuan Wen, Hehe Fan |
AAAI | 2 |
| 2024 | SRL-ProtoNet: Self-supervised representation learning for few-shot remote sensing scene classificationabstractAbstract Using a deep learning method to classify a large amount of labelled remote sensing scene data produces good performance. However, it is challenging for deep learning based methods to generalise to classification tasks with limited data. Few‐shot learning allows neural networks to classify unseen categories when confronted with a handful of labelled data. Currently, episodic tasks based on meta‐learning can effectively complete few‐shot classification, and training an encoder that can conduct representation learning has become an important component of few‐shot learning. An end‐to‐end few‐shot remote sensing scene classification model based on ProtoNet and self‐supervised learning is proposed. The authors design the Pre‐prototype for a more discrete feature space and better integration with self‐supervised learning, and also propose the ProtoMixer for higher quality prototypes with a global receptive field. The authors’ method outperforms the existing state‐of‐the‐art self‐supervised based methods on three widely used benchmark datasets: UC‐Merced, NWPU‐RESISC45, and AID. Compare with previous state‐of‐the‐art performance. For the one‐shot setting, this method improves by 1.21%, 2.36%, and 0.84% in AID, UC‐Merced, and NWPU‐RESISC45, respectively. For the five‐shot setting, this method surpasses by 0.85%, 2.79%, and 0.74% in the AID, UC‐Merced, and NWPU‐RESISC45, respectively. Yansheng Gao, Jianrong Zhang |
IET Comput. Vis. | 5 |
| 2023 | Generating Human Motion from Textual Descriptions with Discrete RepresentationsabstractIn this work, we investigate a simple and must-known conditional generative framework based on Vector Quantised-Variational AutoEncoder (VQ-VAE) and Generative Pre-trained Transformer (GPT) for human motion generation from textural descriptions. We show that a simple CNN-based VQ-VAE with commonly used training recipes (EMA and Code Reset) allows us to obtain high-quality discrete representations. For GPT, we incorporate a simple corruption strategy during the training to alleviate training-testing discrepancy. Despite its simplicity, our T2M-GPT shows better performance than competitive approaches, including recent diffusion-based approaches. For example, on HumanML3D, which is currently the largest dataset, we achieve comparable performance on the consistency between text and generated motion (R-Precision), but with FID 0.116 largely outperforming MotionDiffuse of 0.630. Additionally, we conduct analyses on HumanML3D and observe that the dataset size is a limitation of our approach. Our work suggests that VQ-VAE still remains a competitive approach for human motion generation. Our implementation is available on the project page: https://mael-zys.github.io/T2M-GPT/. Jianrong Zhang, Yangsong Zhang 0002, Xiaodong Cun, Yong Zhang 0034, Hongtao Lu 0001, Xi Shen 0001, Shan Ying |
CVPR | 1 |
| 2023 | Decoupling with Entropy-based Equalization for Semi-Supervised Semantic SegmentationabstractSemi-supervised semantic segmentation methods are the main solution to alleviate the problem of high annotation consumption in semantic segmentation. However, the class imbalance problem makes the model favor the head classes with sufficient training samples, resulting in poor performance of the tail classes. To address this issue, we propose a Decoupled Semi-Supervise Semantic Segmentation (DeS4) framework based on the teacher-student model. Specifically, we first propose a decoupling training strategy to split the training of the encoder and segmentation decoder, aiming at a balanced decoder. Then, a non-learnable prototype-based segmentation head is proposed to regularize the category representation distribution consistency and perform a better connection between the teacher model and the student model. Furthermore, a Multi-Entropy Sampling (MES) strategy is proposed to collect pixel representation for updating the shared prototype to get a class-unbiased head. We conduct extensive experiments of the proposed DeS4 on two challenging benchmarks (PASCAL VOC 2012 and Cityscapes) and achieve remarkable improvements over the previous state-of-the-art methods. Chuanghao Ding, Jianrong Zhang, Henghui Ding, Tengfei Xing, Runbo Hu |
IJCAI | 2 |
| 2022 | Region-level Contrastive and Consistency Learning for Semi-Supervised Semantic SegmentationabstractCurrent semi-supervised semantic segmentation methods mainly focus on designing pixel-level consistency and contrastive regularization. However, pixel-level regularization is sensitive to noise from pixels with incorrect predictions, and pixel-level contrastive regularization has a large memory and computational cost. To address the issues, we propose a novel region-level contrastive and consistency learning framework (RC^2L) for semi-supervised semantic segmentation. Specifically, we first propose a Region Mask Contrastive (RMC) loss and a Region Feature Contrastive (RFC) loss to accomplish region-level contrastive property. Furthermore, Region Class Consistency (RCC) loss and Semantic Mask Consistency (SMC) loss are proposed for achieving region-level consistency. Based on the proposed region-level contrastive and consistency regularization, we develop a region-level contrastive and consistency learning framework (RC^2L) for semi-supervised semantic segmentation, and evaluate our RC^2L on two challenging benchmarks (PASCAL VOC 2012 and Cityscapes), outperforming the state-of-the-art. Jianrong Zhang, Chuanghao Ding, Guodong Guo |
IJCAI | 1 |
| 2017 | Human behaviors modeling in multi-agent virtual environment
Linqin Cai, Jimin Yu, Jianrong Zhang |
Multim. Tools Appl. | 4 |