VLDB 2026 Research / reviewers in the wild / expert
Jingru Tan
dblp:254/1247
· DBLP profile ↗
16ranked-venue papers
5as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 5 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Charts to Code: A Hierarchical Benchmark for Multimodal ModelsabstractJiahao Tang, Henry Hengyuan Zhao, Lijian Wu, Zijian Zhang, Yifei Tao, Dongxing Mao, Yang Wan, Jingru Tan, Min Zeng, Min Li, Alex Jinpeng Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jiahao Tang, Hengyuan Zhao, Lijian Wu, Yifei Tao, Dongxing Mao, Yang Wan, Jingru Tan, Min Zeng 0004, Min Li 0007, Alex Jinpeng Wang |
ACL (1) | 8 |
| 2026 | Poll-Encode-Control: A Reliability-Aware Framework for UAV-Assisted Agricultural IoT Data CollectionabstractEfficient and reliable data collection is essential for agricultural Internet of Things (IoT) systems, where timely sensing supports precision farming. Unmanned aerial vehicle (UAV)-assisted collection provides flexible and low cost coverage. However, in wide and remote fields, UAV-assisted collection faces two coupled challenges. First, agricultural field operations and field-state changes, such as irrigation, rainfall, harvesting, canopy occlusion, machinery movement, and fluctuations in solar exposure, can jointly and unevenly perturb traffic, communication, and energy processes, leading to bursty arrivals, degraded air-to-ground (A2G) links, and variations in harvested energy. Second, under single-link communication, the UAV can directly access only one node in each slot, making frequent global state refresh difficult and causing stale beliefs to accumulate rapidly after abrupt changes. These effects lead to biased scheduling, unnecessary maneuvering, packet loss, and excess energy consumption. To address this problem, this paper proposes a three stage framework for joint node scheduling and UAV trajectory control. It first refreshes selected node states to correct stale beliefs and then encodes refreshed and inferred states through a reliability-aware Transformer. Finally, it performs conservative hybrid-action control to coordinate discrete scheduling and continuous motion while mitigating value overestimation under partial observability and non-stationarity. Simulations under point-shift, cluster-shift, and global-shift scenarios show that the proposed method consistently reduces packet loss, improves energy efficiency, and achieves faster recovery than representative reinforcement learning (RL) and heuristic baselines. Jingru Tan, Tom H. Luan, Wenbo Guan, Jinkai Zheng |
IEEE Internet Things J. | 1 |
| 2025 | OmniBal: Towards Fast Instruction-Tuning for Vision-Language Models via Omniverse Computation BalanceabstractVision-language instruction-tuning models have recently achieved significant performance improvements. In this work, we discover that large-scale 3D parallel training on those models leads to an imbalanced computation load across different devices. The vision and language parts are inherently heterogeneous: their data distribution and model architecture differ significantly, which affects distributed training efficiency. To address this issue, we rebalance the computational load from data, model, and memory perspectives, achieving more balanced computation across devices. Specifically, for the data, instances are grouped into new balanced mini-batches within and across devices. A search-based method is employed for the model to achieve a more balanced partitioning. For memory optimization, we adaptively adjust the re-computation strategy for each partition to utilize the available memory fully. These three perspectives are not independent but are closely connected, forming an omniverse balanced training framework. Extensive experiments are conducted to validate the effectiveness of our method. Compared with the open-source training code of InternVL-Chat, training time is reduced greatly, achieving about 1.8$\times$ speed-up. Our method’s efficacy and generalizability are further validated across various models and datasets. Codes will be released at https://github.com/ModelTC/OmniBal. Yongqiang Yao, Jingru Tan, Feizhao Zhang, Yazhe Niu, Xin Jin 0008, Bo Li 0126, Pengfei Liu 0003, Ruihao Gong, Dahua Lin, Ningyi Xu |
ICML | 2 |
| 2025 | Hierachical Balance Packing: Towards Efficient Supervised Fine-tuning for Long-Context LLMabstractTraining Long-Context Large Language Models (LLMs) is challenging, as hybrid training with long-context and short-context data often leads to workload imbalances. Existing works mainly use data packing to alleviate this issue, but fail to consider imbalanced attention computation and wasted communication overhead. This paper proposes Hierarchical Balance Packing (HBP), which designs a novel batch-construction method and training recipe to address those inefficiencies. In particular, the HBP constructs multi-level data packing groups, each optimized with a distinct packing length. It assigns training samples to their optimal groups and configures each group with the most effective settings, including sequential parallelism degree and gradient checkpointing configuration. To effectively utilize multi-level groups of data, we design a dynamic training pipeline specifically tailored to HBP, including curriculum learning, adaptive sequential parallelism, and stable loss. Our extensive experiments demonstrate that our method significantly reduces training time over multiple datasets and open-source models while maintaining strong performance. For the largest DeepSeek-V2 (236B) MoE model, our method speeds up the training by 2.4$\times$ with competitive performance. Codes will be released at https://github.com/ModelTC/HBP. Yongqiang Yao, Jingru Tan, Kaihuan Liang, Feizhao Zhang, Yazhe Niu, Ruihao Gong, Dahua Lin, Ningyi Xu |
NeurIPS | 2 |
| 2025 | Robust long-tailed recognition with distribution-aware adversarial example generation
Bo Li 0126, Yongqiang Yao, Jingru Tan, Dandan Zhu 0001, Ruihao Gong, Ye Luo 0004 |
Neural Networks | 3 |
| 2025 | LSTM-Characterized Approach for Chip Floorplanning: Leveraging HyperGCN and DRQNabstractIn the field of very large-scale integration (VLSI) chip design, chip floorplanning plays a crucial role as it directly influences key optimization objectives such as placement wirelength. This, in turn, affects signal delay, power efficiency, routability, and overall cost. However, traditional reinforcement learning (RL) methods for chip floorplanning often oversimplify this complex and dynamic task. They tend to overlook the cascading effects of module placements and fail to fully comprehend the intricate interdependencies that are vital for making informed decisions. To address these challenges, we introduce an innovative approach that combines hypergraph graph convolutional networks (HyperGCNs) with deep recurrent Q-networks (DRQNs). This integration allows us to capture the nuanced dynamics and interconnected aspects of chip design more effectively. We enhance the traditional Markov decision process (MDP) model by incorporating a state characterization layer based on long short-term memory (LSTM) technology. Initially, HyperGCN efficiently encodes netlist information, simplifying complex graph structures into lower dimensional vectors, thereby enhancing knowledge processing. Subsequently, we treat the chip as an agent and apply the DRQN algorithm to optimize the module layout. DRQN’s LSTM utilizes a recurrent layer structure to grasp dependencies between modules, combined with the deep Q-network (DQN)’s optimization capabilities, enabling us to navigate the complexities of floorplanning. Our approach improves the state representation by encompassing a broader understanding of interconnected module characteristics. This allows our agents to make decisions that take into account the collective impact of module adjustments, rather than viewing each change in isolation. This comprehensive state representation, which includes diverse features and their evolving relationships, significantly enhances the decision-making capabilities of our agents. Our extensive experiments demonstrate that our method outperforms traditional heuristic-based and other learning-based floorplanning techniques. To the best of our knowledge, this is the first application of LSTM with DQN in circuit design, representing a significant advancement in the field. Wenbo Guan, Xiaoyan Tang, Jingru Tan, Yimen Zhang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2024 | From Isolated Islands to Pangea: Unifying Semantic Space for Human Action UnderstandingabstractAction understanding has attracted long-term attention. It can be formed as the mapping from the physical space to the semantic space. Typically, researchers built datasets according to idiosyncratic choices to define classes and push the envelope of benchmarks respectively. Datasets are incompatible with each other like “Isolated Islands” due to semantic gaps and various class granularities, e.g., do housework in dataset A and wash plate in dataset B. We argue that we need a more principled semantic space to concentrate the community efforts and use all datasets together to pursue generalizable action learning. To this end, we design a structured action semantic space in view of verb taxonomy hierarchy and covering massive actions. By aligning the classes of previous datasets to our semantic space, we gather (image/video/skeleton/McCap] datasets into a unified database in a unified label system, i.e., bridging “isolated islands” into a “Pangea”. Accordingly, we propose a novel model mapping from the physical space to semantic space to fully use Pangea. In extensive experiments, our new system shows significant superiority, especially in transfer learning. Our code and data will be made public at https://mvig-rhos.com/pangea. Yong-Lu Li 0001, Xinpeng Liu 0002, Yiming Dou, Yikun Ji, Junyi Zhang 0004, Yixing Li, Jingru Tan, Cewu Lu |
CVPR | 10 |
| 2024 | Transformer-Characterized Approach for Chip Floorplanning: Leveraging HyperGCN and DTQNabstractIn the realm of very large-scale integration (VLSI) chip design, chip floorplanning is essential, directly impacting key optimization objectives like placement wirelength, which in turn affects signal delay, power efficiency, routability, and overall cost. Traditional reinforcement learning (RL) methods for chip floorplanning often oversimplify the complex, dynamic nature of this task, typically overlooking the cascading effects of module placements and failing to fully grasp the intricate interdependencies crucial for informed decision-making. To address these challenges, we introduce an innovative approach that fuses hypergraph graph convolutional networks (HyperGCN) with deep Transformer Q-networks (DTQN). This integration captures the nuanced dynamics and interconnected aspects of chip design more effectively. We enhance the traditional Markov decision process (MDP) model with a state characterization layer based on Transformer technology. Initially, HyperGCN effectively encodes netlist information, simplifying complex graph structures into lower-dimensional vectors, thus enhancing knowledge processing. Subsequently, viewing the chip as an agent, we apply the DTQN algorithm to optimize the module layout. DTQN's Transformer encoder utilizes a multi-head self-attention (MSA) mechanism to grasp long-range dependencies between modules, coupled with DQN's optimization capabilities, to navigate the complexities of floorplanning. Our approach enhances the state representation by incorporating a broader understanding of the interconnected module characteristics, which allows agents to make decisions that account for the collective impact of module adjustments, rather than viewing each change in isolation. This comprehensive state representation, encompassing diverse features and their evolving relationships, significantly enhances the agent's decision-making capabilities. Our extensive experiments demonstrate that our method outperforms traditional heuristic-based and other learning-based floorplanning techniques. To our knowledge, this is the first application of the Transformer in circuit design, marking a significant advancement in the field. Wenbo Guan, Xiaoyan Tang, Jingru Tan |
ICCD | 4 |
| 2024 | A Data Synchronization Incentive Scheme in Vehicular Digital Twin Network with Stackelberg GameabstractThe evolving digital twin technology translates physical entities into the digital realm, allowing the exploration of abundant digital resources to optimize the task execution of these physical entities. Real-time data synchronization between physical entities and their digital twins is essential for the effective functioning of digital twin systems. In this paper, we investigate the challenge of data synchronization in vehicular digital twin networks operating in open street scenarios, where multiple vehicles rely on cellular networks for continuous data synchronization with their digital twins. Given the contention for cellular bandwidth among vehicles, a coordination scheme is required to manage resource allocation. As vehicles are fully distributed driven by self-interests only, a game-theoretic approach is proposed that leverages a cloud center controller to guide the sharing of cellular resources among digital twins. An optimal incentive mechanism is introduced to encourage digital twins to adhere to the center's guidance, promoting global social welfare. Through extensive simulations, we demonstrate that the proposed scheme successfully motivates vehicles to follow the center's guidance, leading to efficient data synchronization and mutual benefit maximization. Jingru Tan, Jinkai Zheng, Tom H. Luan, Longxiang Gao, Zhou Su 0001 |
VTC Spring | 1 |
| 2024 | Rectify representation bias in vision-language models for long-tailed recognition
Bo Li 0126, Yongqiang Yao, Jingru Tan, Ruihao Gong, Ye Luo 0004 |
Neural Networks | 3 |
| 2023 | The Equalization Losses: Gradient-Driven Training for Long-tailed Object RecognitionabstractLong-tail distribution is widely spread in real-world applications. Due to the extremely small ratio of instances, tail categories often show inferior accuracy. In this paper, we find such performance bottleneck is mainly caused by the imbalanced gradients, which can be categorized into two parts: (1) positive part, deriving from the samples of the same category, and (2) negative part, contributed by other categories. Based on comprehensive experiments, it is also observed that the gradient ratio of accumulated positives to negatives is a good indicator to measure how balanced a category is trained. Inspired by this, we come up with a gradient-driven training mechanism to tackle the long-tail problem: re-balancing the positive/negative gradients dynamically according to current accumulative gradients, with a unified goal of achieving balance gradient ratios. Taking advantage of the simple and flexible gradient mechanism, we introduce a new family of gradient-driven loss functions, namely equalization losses. We conduct extensive experiments on a wide spectrum of visual tasks, including two-stage/single-stage long-tailed object detection (LVIS), long-tailed image classification (ImageNet-LT, Places-LT, iNaturalist), and long-tailed semantic segmentation (ADE20 K). Our method consistently outperforms the baseline models, demonstrating the effectiveness and generalization ability of the proposed equalization losses. Jingru Tan, Bo Li 0126, Yongqiang Yao, Fengwei Yu, Tong He 0001, Wanli Ouyang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Equalized Focal Loss for Dense Long-Tailed Object DetectionabstractDespite the recent success of long-tailed object detection, almost all long-tailed object detectors are developed based on the two-stage paradigm. In practice, one-stage detectors are more prevalent in the industry because they have a simple and fast pipeline that is easy to deploy. However, in the long-tailed scenario, this line of work has not been explored so far. In this paper, we investigate whether one-stage detectors can perform well in this case. We discover the primary obstacle that prevents one-stage detectors from achieving excellent performance is: categories suffer from different degrees of positive-negative imbalance problems under the long-tailed data distribution. The conventional focal loss balances the training process with the same modulating factor for all categories, thus failing to handle the long-tailed problem. To address this issue, we propose the Equalized Focal Loss (EFL) that rebalances the loss contribution of positive and negative samples of different categories independently according to their imbalance degrees. Specifically, EFL adopts a category-relevant modulating factor which can be adjusted dynamically by the training status of different categories. Extensive experiments conducted on the challenging LVIS v1 benchmark demonstrate the effectiveness of our proposed method. With an end-to-end training pipeline, EFL achieves 29.2% in terms of overall AP and obtains significant performance improvements on rare categories, surpassing all existing state-of-the-art methods. The code is available at https: //github.com/ModelTC/EOD. Bo Li 0126, Yongqiang Yao, Jingru Tan, Fengwei Yu, Ye Luo 0004 |
CVPR | 3 |
| 2021 | Equalization Loss v2: A New Gradient Balance Approach for Long-Tailed Object DetectionabstractRecently proposed decoupled training methods emerge as a dominant paradigm for long-tailed object detection. But they require an extra fine-tuning stage, and the dis-jointed optimization of representation and classifier might lead to suboptimal results. However, end-to-end training methods, like equalization loss (EQL), still perform worse than decoupled training methods. In this paper, we re-veal the main issue in long-tailed object detection is the imbalanced gradients between positives and negatives, and find that EQL does not solve it well. To address the problem of imbalanced gradients, we introduce a new version of equalization loss, called equalization loss v2 (EQL v2), a novel gradient guided reweighing mechanism that re-balances the training process for each category independently and equally. Extensive experiments are performed on the challenging LVIS benchmark. EQL v2 outperforms origin EQL by about 4 points overall AP with 14 ∼ 18 points improvements on the rare categories. More importantly, it also surpasses decoupled training methods. With-out further tuning for the Open Images dataset, EQL v2 improves EQL by 7.3 points AP, showing strong generalization ability. Codes have been released at https://github.com/tztztztztz/eqlv2 Jingru Tan, Xin Lu 0002, Changqing Yin, Quanquan Li |
CVPR | 1 |
| 2021 | RefineMask: Towards High-Quality Instance Segmentation With Fine-Grained FeaturesabstractThe two-stage methods for instance segmentation, e.g. Mask R-CNN, have achieved excellent performance recently. However, the segmented masks are still very coarse due to the downsampling operations in both the feature pyramid and the instance-wise pooling process, especially for large objects. In this work, we propose a new method called RefineMask for high-quality instance segmentation of objects and scenes, which incorporates fine-grained features during the instance-wise segmenting process in a multi-stage manner. Through fusing more detailed information stage by stage, RefineMask is able to refine high-quality masks consistently. RefineMask succeeds in segmenting hard cases such as bent parts of objects that are oversmoothed by most previous methods and outputs accurate boundaries. Without bells and whistles, RefineMask yields significant gains of 2.6, 3.4, 3.8 AP over Mask R-CNN on COCO, LVIS, and Cityscapes benchmarks respectively at a small amount of additional computational cost. Furthermore, our single-model result outperforms the winner of the LVIS Challenge 2020 by 1.3 points on the LVIS test-dev set and establishes a new state-of-the-art. Code will be available at https://github.com/zhanggang001/RefineMask. Xin Lu 0002, Jingru Tan, Jianmin Li 0001, Zhaoxiang Zhang 0001, Quanquan Li, Xiaolin Hu 0001 |
CVPR | 3 |
| 2021 | Liver Tumor Detection Via A Multi-Scale Intermediate Multi-Modal Fusion Network on MRI ImagesabstractAutomatic liver tumor detection can assist doctors to make effective treatments. However, how to utilize multi-modal images to improve detection performance is still challenging. Common solutions for using multi-modal images consist of early, inter-layer, and late fusion. They either do not fully consider the intermediate multi-modal feature interaction or have not put their focus on tumor detection. In this paper, we propose a novel multi-scale intermediate multi-modal fusion detection framework to achieve multi-modal liver tumor detection. Unlike early or late fusion, it maintains two branches of different modal information and introduces cross-modal feature interaction progressively, thus better leveraging the complementary information contained in multi-modalities. To further enhance the multi-modal context at all scales, we design a multi-modal enhanced feature pyramid. Extensive experiments on the collected liver tumor magnetic resonance imaging (MRI) dataset show that our framework outperforms other state-of-the-art detection approaches in the case of using multi-modal images. Peiyun Zhou, Jingru Tan, Baoye Sun, Ruoyu Guan, Zhutao Wang, Ye Luo 0004 |
ICIP | 3 |
| 2020 | Equalization Loss for Long-Tailed Object RecognitionabstractObject recognition techniques using convolutional neural networks (CNN) have achieved great success. However, state-of-the-art object detection methods still perform poorly on large vocabulary and long-tailed datasets, e.g. LVIS. In this work, we analyze this problem from a novel perspective: each positive sample of one category can be seen as a negative sample for other categories, making the tail categories receive more discouraging gradients. Based on it, we propose a simple but effective loss, named equalization loss, to tackle the problem of long-tailed rare categories by simply ignoring those gradients for rare categories. The equalization loss protects the learning of rare categories from being at a disadvantage during the network parameter updating. Thus the model is capable of learning better discriminative features for objects of rare classes. Without any bells and whistles, our method achieves AP gains of 4.1% and 4.8% for the rare and common categories on the challenging LVIS benchmark, compared to the Mask R-CNN baseline. With the utilization of the effective equalization loss, we finally won the 1st place in the LVIS Challenge 2019. Code has been made available at: https://github.com/tztztztztz/eql.detectron2. Jingru Tan, Changbao Wang, Buyu Li, Quanquan Li, Wanli Ouyang, Changqing Yin |
CVPR | 1 |