VLDB 2026 Research / reviewers in the wild / expert
RenGang Li
dblp:262/1327 · also Rengang Li
· DBLP profile ↗
44ranked-venue papers
5as first author
42since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 2 first-author · 14 since 2021Systems, architecture and hardware · 6 · 1 first-author · 6 since 2021Computer networks · 6 · 2 first-author · 6 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Theory of computation · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | T-MSA: Transformer-Driven Multi-Strategy Adaptive Microarchitecture Design Space ExplorationabstractThe design of modern processors ignores the topological relationships among all design parameters, leading to significant simulation costs wasted on invalid designs. Therefore, we propose the T-MSA to address this issue. It is a Transformer-driven multi-strategy adaptive design space exploration scheme. A customized lightweight Transformer (LiteFormer) is devised to model topological relationships among arbitrary design parameters, constructing an implicit interaction graph in the latent space. Secondly, we design a dynamic active learning (DynamicAL) strategy to extract sparse and high-quality initial points via sparse centroid initialization and hybrid sampling. Finally, a triple Pareto frontier acquisition function (TriPFAF) is devised to guide optimization direction based on gains from three types of Pareto frontiers, dynamically balancing exploration and exploitation. We conducted rigorous experiments on two BOOM evaluation platforms, demonstrating that T-MSA efficiently and comprehensively optimizes the performance-power-area (PPA) objective. The designs it identifies achieve significant improvements over state-of-the-art DSE algorithms on Pareto hypervolume (HV). When attaining the same HV value, T-MSA outperforms BOOM-Explorer by 188.24% and 133.33% on two platforms. Fan Yang 0032, Xiaochuan Li 0001, Cong Xu 0001, RenGang Li, Baoyu Fan |
DATE | 7 |
| 2025 | ATQ: An Optimal Gradient Quantization Strategy for Distributed Training Under Relatively Low-Bandwidth and Unstable NetworkabstractLarge-scale distributed computing environments have become an indispensable foundation for training Large Language Models (LLMs). Although many data centers now have dedicated networks with higher bandwidth, the network remains a bottleneck for distributed training compared to high-speed dedicated connections between GPUs. Gradient compression lies in training LLMs using common computing resources, often with relatively lower bandwidth and unstable networks, particularly as the demand for democratized AI continues to rise. In such scenarios, the training time cost disproportionately increases as network quality degrades. To strike a balance between model performance and total training time, we propose the Automatic Transmission Quantization (ATQ) strategy. Considering factors such as bandwidth, training loss, and training time, the ATQ strategy dynamically adjusts the gradient quantization scale throughout the training process. In experiments, our ATQ method has been shown to reduce the total training time by$\mathbf{5 7. 5 \%}$compared to baseline. In real-world training tasks under low-quality network conditions, the time savings effect achieved by ATQ is expected to be even more significant. RenGang Li, Yujie Dai, Kefeng Zhu, Pengyu Hou |
HPCC | 1 |
| 2025 | MMEditor: Multimodal Prompt-Driven 3D Gaussian Splatting EditingabstractWe propose a multimodal 3D scene editing framework MMEditor to create or modify objects within an extant 3D Gaussian Splatting (3DGS) according to text and image prompts. MMEditor employs a multimodal image editing module to iteratively optimize 3D Gaussians in editing regions for delicate and multi-view consistent 3D editing. The key multimodal image editing module can perform editing with accurate appearance and location control, which is achieved by two designs. First, a multimodel adapter block takes the reference image as a foreign language to augment the text prompt, enabling editing results to align with the generic text description and the unique characteristics in the reference image. Second, an attention-based localization block localizes cross-attention with user-defined 3D bounding boxes, thereby ensuring the editing occurs in editing regions. Experiments demonstrate that our method achieves more accurate and controllable results than previous state-of-the-art methods. RenGang Li, Yaqian Zhao, Xiaohui Zhang 0017, Hui Wei 0005, Ruyang Li |
ICASSP | 2 |
| 2025 | Improving Height Prediction for Vision-Based Roadside 3D Object DetectionabstractRoadside vision-based 3D object detection is vital in many applications, such as autonomous driving. The mainstream methods enhance the accuracy of distance estimation by converting predicted height distribution into depth distribution. However, predicting object’s height in roadside perception is challenging, particularly for distant and small objects. Therefore, this work proposes a series of methods to optimize height prediction. Firstly, we propose depth and height decoding supervision methods to optimize the height network by supervising the decoded depth and height values. Then, a height distribution alignment loss is introduced to optimize the height network by constraining the consistency of height distributions among objects of the same category. Experimental results demonstrate that the proposed optimization methods can effectively improve the accuracy of roadside vision-based 3D object detection. For instance, our methods improve the accuracies by 2.31%, 4.33%, and 4.24% for the cyclist category at easy, medium, and hard levels, respectively. Tengfei Zhang 0004, RenGang Li, Yaqian Zhao, Ruyang Li |
ICASSP | 3 |
| 2025 | Dropletvideo: A Dataset and Approach to Explore Integral Spatio-Temporal Consistent Video Generation
Guoguang Du 0001, Xiaochuan Li 0001, Qi Jia 0004, Lu Liu 0009, Cong Xu 0001, Zhenhua Guo 0003, Yaqian Zhao, Xiaoli Gong, RenGang Li, Baoyu Fan |
ICCV | 12 |
| 2025 | Proactive Fault-tolerance Driven Task Scheduling System for IoV Edge NetworksabstractThe emergence of Internet of Vehicles (IoV) technology provides a wider range of application scenarios for edge computing based on Vehicle-to-everything (V2X). It is essential to ensure the high availability and reliability of services in IoV systems. Currently, cloud service providers have established a data center level of fault tolerance, such as redundancy and checkpoints, guaranteeing the reliability of cloud infrastructure and reducing phenomena such as service termination or downtime. However, current computing systems reactively handle failures. Especially in edge computing, this approach not only lacks flexibility but also consumes excessive system resources, which is not conducive to ensuring the reliability in resource-constrained systems and poses security risks to end users. To mitigate this problem, we propose a Proactive Fault-tolerance Driven Task Scheduling System. Different from the traditional reactive strategies, the proposed framework predicts the possible system crashes by monitoring the critical state indicators of the computing system. According to the prediction results, a class of tasks or services that are most likely to be terminated are rescheduled in advance. Extensive experiments are conducted, and evaluation results demonstrate that our proposed proactive fault tolerance framework can effectively improve the long-term performance of the IoV edge system. Yaqiang Zhang, RenGang Li, Yaqian Zhao, Hongzhi Shi, Guangyuan Xu |
ICNP | 2 |
| 2025 | Asymptotically Optimal Repair of Reed-Solomon Codes with Small Sub-Packetization under Rack-Aware ModelabstractThis paper presents a comprehensive study on the asymptotically optimal repair of Reed-Solomon (RS) codes with small sub-packetization, specifically tailored for rack-aware distributed storage systems. Through the utilization of multibase expansion, we introduce a novel approach that leverages monomials to construct linear repair schemes for RS codes. Our repair schemes which adapt to all admissible parameters achieve asymptotically optimal repair bandwidth while significantly reducing the sub-packetization compared with existing schemes. Furthermore, our approach is capable of repairing RS codes with asymptotically optimal repair bandwidth under the homogeneous storage model, achieving smaller sub-packetization than existing methods. Zhongyan Liu, RenGang Li, Yaqian Zhao, Yaqiang Zhang |
ITW | 3 |
| 2025 | Tripartite interaction representation learning for multi-modal sentiment analysis
Binqiang Wang, Gang Dong, Yaqian Zhao, RenGang Li, Wenfeng Yin |
Expert Syst. Appl. | 4 |
| 2025 | SFP: Similarity-based filter pruning for deep neural networks
RenGang Li, Chaoyao Shen, Xiaofeng Zou, Jiuyang Wang, Nanjun Li |
Inf. Sci. | 2 |
| 2025 | Algorithm-Hardware Co-design for Accelerating Depthwise Separable CNNsabstractDepthwise separable convolution (DSC) is a popular method for constructing lightweight neural networks. However, the pointwise convolution (PWC) has a much larger number of parameters than the depthwise convolution (DWC), causing the imbalanced parameter ratio of PWC to DWC. In this article, we propose an efficient and hardware-efficiency convolution (Shared Kernel sliding on channel Convolution, SKC) to replace the redundant PWC in DSC for a balanced parameter ratio, where SKC customizes the sharing kernel in the channel dimension to reduce the number of parameters, and the local connection in the channel dimension reduces the computation. Furthermore, the proposed SKC is suitable for Winograd acceleration, and the large kernel decomposition method is introduced to facilitate its use. We implement the first Winograd-based FPGA hardware accelerator for DSCNets. The shared 1D and 2D Winograd convolution computing engine is proposed to compute the proposed DSC consisting of DWC and SKC efficiently. An alternating loading and reusing storage approach is developed to efficiently load SKC input feature maps. Experimental results show our DSC-based accelerator can achieve 20× higher power efficiency at the cost of a small loss of accuracy by algorithm-hardware co-design compared with traditional accelerators. RenGang Li, Tinghuan Chen, Meng Zhang 0010, Henk Corporaal |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2025 | FrigateBird: Decoupling Metadata/Data Services for Continuously Fast Object StorageabstractCurrently, ride-hailing service has tens of millions of registered drivers and hundreds of millions of registered passengers, serving tens of millions of rides per day. Different from social media applications like Facebook and LinkedIn, ride hailing needs to read/write/query a large number of small objects (like photos and audio/video pieces)always fast, so that it can support critical online computations such as face comparison and sentiment analysis on audio/video records. This is of particular importance for ride-hailing service to recognize and avoid potential dangers. Existing object stores (like Haystack and Tectonic) usually store object data in files and place object metadata in a separate key-value store (like RocksDB), which is unsuitable for ride-hailing service mainly because the objects' file-related information is placed together with the object data. This severely affects the I/O performance of object storage: first, for crash consistency, the writes of object data and object metadata must be conducted inseparatephases of one transaction, which significantly increases I/O latency; second, the space of deleted objects needs to be reclaimed viacompaction, which could sharply lower I/O performance when the system is busy in serving normal read/write requests. This paper describes FrigateBird, a continuously fast object store for ride-hailing service. FrigateBird differs from existing object stores in three aspects. First, we present a metadata/data decoupled service architecture for object storage, where therichmetadata service realizes efficient queries and updates of object metadata, and therawdata service purely performs disk I/O to read/write object data from/to raw disks. Second, we propose a rich metadata structure (calledRichmeta) taking the write operation logs as part of object metadata, which allows FrigateBird tosimultaneouslywrite the object data to raw disks (without filesystem overhead) and write the metadata to a key-value store, guaranteeing crash consistency by checking whether the transaction is completed and rolling back if not. Third, we design a compaction-free deletion mechanism which can efficiently delete an object by only updating the metadata without involving the data service, so that FrigateBird can efficiently support ride-hailing service's frequent delete operations while avoiding data-migration-caused performance hiccup. Evaluation shows that FrigateBird outperforms the state-of-the-art object stores by up to$3.36\times$and$27.9\times$in the mean I/O latency for normal and in-compaction scenarios, respectively. Yiming Zhang 0003, Ke-Kun Hu, Gang Dong, RenGang Li |
IEEE Trans. Serv. Comput. | 5 |
| 2024 | Image Content Generation with Causal ReasoningabstractThe emergence of ChatGPT has once again sparked research in generative artificial intelligence (GAI). While people have been amazed by the generated results, they have also noticed the reasoning potential reflected in the generated textual content. However, this current ability for causal reasoning is primarily limited to the domain of language generation, such as in models like GPT-3. In visual modality, there is currently no equivalent research. Considering causal reasoning in visual content generation is significant. This is because visual information contains infinite granularity. Particularly, images can provide more intuitive and specific demonstrations for certain reasoning tasks, especially when compared to coarse-grained text. Hence, we propose a new image generation task called visual question answering with image (VQAI) and establish a dataset of the same name based on the classic Tom and Jerry animated series. Additionally, we develop a new paradigm for image generation to tackle the challenges of this task. Finally, we perform extensive experiments and analyses, including visualizations of the generated content and discussions on the potentials and limitations. The code and data are publicly available under the license of CC BY-NC-SA 4.0 for academic and non-commercial usage at: https://github.com/IEIT-AGI/MIX-Shannon/blob/main/projects/VQAI/lgd_vqai.md. Xiaochuan Li 0001, Baoyu Fan, Zhenhua Guo 0003, Yaqian Zhao, RenGang Li |
AAAI | 8 |
| 2024 | A Distributed Algorithm for Rumor Blocking on Social Networks
Ruidong Yan, Zhenhua Guo 0003, Yaqian Zhao, RenGang Li, Xingjian Ding |
COCOON (2) | 4 |
| 2024 | CSOT: Cross-scan Object Transfer for Semi-Supervised LiDAR Object Detection
Jinglin Zhan, Tiejun Liu, RenGang Li, Zhaoxiang Zhang 0001, Yuntao Chen |
ECCV (17) | 3 |
| 2024 | DM-SARAH: A Variance Reduction Optimization Algorithm for Machine Learning SystemsabstractNowadays, the variance reduction (VR) technique is used to improve the performance of gradient-type algorithms in machine learning and deep learning. However, some existing VR algorithms require unrealistic assumptions or conditions such as τ-gradient dominated and Polyak-Lojasiewicz (PL) conditions, which limit their applications. In this paper, we present a Double Mini-batch StochAstic Recursive grAdient algoritHm (DM-SARAH) without these assumptions or conditions to solve the convex and non-convex optimization problems respectively. The main contributions of this paper are twofold: (1) At the theoretical level, we optimize the convergence rate and provide a complexity analysis of DM-SARAH, and (2) At the experimental level, we evaluate the effectiveness and efficiency of the proposed algorithm on various datasets. The experimental results indicate that the proposed algorithm outperforms existing methods. RenGang Li, Ruidong Yan, Zhenhua Guo 0003, Zhi-Yong Qiu, Yaqian Zhao |
GLOBECOM | 1 |
| 2024 | FHNTT: a flexible Number Theoretic Transform design based on hybrid-radix butterflyabstractEmerging technologies, such as cloud computing and artificial intelligence, significantly arouse concern about data security and privacy. Homomorphic encryption (HE) is a promising invention, which enables computation on encrypted data without decrypting it so as to ensure data security and privacy. Nevertheless, computation within homomorphic encryption involves time-consuming operations, e.g., Number Theoretic Transform (NTT). The tremendous computation overhead is the critical obstacle in deploying HE applications widely. Besides, in order to meet the performance and security requirements of different applications, it is pivotal to design parametric NTT architecture. In this paper, we propose a flexible and parametric NTT accelerating scheme based on hybrid-radix butterfly, named FHNTT. Specifically, we construct high radix butterfly units and divide the computation of them into several stages such that every stage can be performed pipelined. The number of required twiddle factors declines with the increase of radix value. In addition, we adopt address offset strategy to reduce memory consumption. We implement FHNTT on FPGA due to its fine-grained parallel computing capabilities and customized architecture. Empirical results show that FHNTT has an improved performance compared with other NTT architectures and supports a wide range of parameters. Concretely, FHNTT achieves up to 1.99 × to 2.78 × improvement in latency over other FPGA implementations and the memory utilization rate is up to 94%. Moreover, the flexibility makes FHNTT applicable to multiple use cases. RenGang Li, Yaqian Zhao, Ruyang Li, Zhiyuan Su, Xuelei Li |
ISPA | 2 |
| 2024 | A Pseudo-Hierarchical Planning Framework with Dynamic-Aware Reinforcement Learning for Autonomous DrivingabstractReinforcement Learning (RL) over motion skill space has been verified to generate more diverse behaviors than that over low-level control space, and has exhibited superior autonomous driving performance in complex traffic scenarios. However, the incomplete observations pose challenges in achieving efficient skill exploration under unsupervised conditions, hampering the driving performance and applicability. In this paper, we propose a dynamic-aware RL with hybrid network (Da-HnRL) to develop a pseudo-hierarchical planning framework for better motion skill learning in challenging dense traffics. Based on the semi-POMDP modeling, we construct a hybrid network with skip connections as the RL backbone, facilitating a better understanding of the underlying system dynamics. Then we design an efficiency-oriented reward shaping mechanism to incentivize active skill exploration, promoting enhanced trade-off between exploration and exploitation. Furthermore, we provide a comprehensive scoring mechanism for policy identification, ensuring the near-optimality. We validate the proposed methods on challenging dense-traffic tasks. The results demonstrate the superiority of our approach over previous methods, with improved learning efficiency, driving stability and generalization. Yaqian Zhao, RenGang Li, Qifu Hu, Tengfei Zhang 0004, Ruyang Li |
IV | 3 |
| 2024 | Infer Induced Sentiment of Comment Response to Video: A New Task, Dataset and BaselineabstractExisting video multi-modal sentiment analysis mainly focuses on the sentiment expression of people within the video, yet often neglects the induced sentiment of viewers while watching the videos. Induced sentiment of viewers is essential for inferring the public response to videos and has broad application in analyzing public societal sentiment, effectiveness of advertising and other areas. The micro videos and the related comments provide a rich application scenario for viewers’ induced sentiment analysis. In light of this, we introduces a novel research task, Multimodal Sentiment Analysis for Comment Response of Video Induced(MSA-CRVI), aims to infer opinions and emotions according to comments response to micro video. Meanwhile, we manually annotate a dataset named Comment Sentiment toward to Micro Video (CSMV) to support this research. It is the largest video multi-modal sentiment dataset in terms of scale and video duration to our knowledge, containing 107, 267 comments and 8, 210 micro videos with a video duration of 68.83 hours. To infer the induced sentiment of comment should leverage the video content, we propose the Video Content-aware Comment Sentiment Analysis (VC-CSA) method as a baseline to address the challenges inherent in this new task. Extensive experiments demonstrate that our method is showing significant improvements over other established baselines. We make the dataset and source code publicly available at https://github.com/IEIT-AGI/MSA-CRVI. Qi Jia 0004, Baoyu Fan, Cong Xu 0001, Lu Liu 0009, Guoguang Du 0001, Zhenhua Guo 0003, Yaqian Zhao, Xuanjing Huang 0001, RenGang Li |
NeurIPS | 10 |
| 2024 | MVIndEmo: a dataset for micro video public-induced emotion prediction on social mediaabstractAbstract Distinct from the realm of perceived emotion research, induced emotion pertains to the emotional responses engendered within content consumers. This facet has garnered considerable attention and finds extensive application in the analysis of public social media. However, the advent of micro videos presents unique challenges when attempting to discern the induced emotional patterns exhibited by content consumers, owing to their free-style representation and other factors. Consequently, we have put forth two novel tasks concerning the recognition of public-induced emotion on micro videos: emotion polarity and emotion classification. Additionally, we have introduced a accessible dataset specifically tailored for the analysis of public-induced emotion on micro videos. The data corpus has been meticulously collected from Tiktok, a burgeoning social media platform renowned for its trendsetting content. To construct the dataset, we have selected eight captivating topics that elicit vibrant social discussions. In devising our label generation strategy, we have employed an automated approach characterized by the fusion of multiple expert models. This strategy incorporates a confidence measure method that relies on three distinct models for effectively aggregating user comments. To accommodate adaptable benchmark configurations, we provide both binary classification labels and probability distribution labels. The dataset encompasses a vast collection of 7,153 labeled micro videos. We have undertaken an extensive statistical analysis of the dataset to provide a comprehensive overview composition. It is our earnest aspiration that this dataset will serve as a catalyst for pioneering research avenues in the analysis of emotional patterns and the understanding of multi-modal information. Zhenhua Guo 0003, Qi Jia 0004, Baoyu Fan, Cong Xu 0001, Yaqian Zhao, RenGang Li |
Multim. Syst. | 8 |
| 2024 | Inexactly Matched Referring Expression Comprehension With RationaleabstractReferring Expression Comprehension (REC) is a multimodal comprehension task that aims to locate an object in an image, given a text description. Traditionally, during the existing REC tasks, there has been a basic assumption that the given text expression and the image are usually exactly matched to each other. However, in real-world scenarios, there is uncertainty in how well the image and text match each other exactly. Illegible objects in the image or ambiguous phrases in the text have the potential to significantly degrade the performance of conventional REC tasks. To overcome these limitations, we consider a more practical and comprehensive REC task, where the given image and its referring text expression can be inexactly matched. Our models aim to correct such inexact matching and supply corresponding interpretations. We refer to this task asFurther REC (FREC). This task is divided into three subtasks: 1) correcting the erroneous text expression using visual information, 2) generating the rationale for this input expression, and 3) localizing the proper object based on the corrected expression. We introduce three new datasets for FREC:Further-RefCOCOs,Further-CopsrefandFurther-Talk2Car. These datasets are based on the existing REC datasets, including RefCOCO and Talk2Car. We developed a novel pipeline architecture to execute the three subtasks simultaneously in an end-to-end fashion. Next, we developed an elastic masked language modeling (EMLM) training head to rectify text errors with uncertain lengths. Our experimental results demonstrate the validity of our proposed pipeline. We hope this work sparks more research focused on inexactly matched REC. Xiaochuan Li 0001, Baoyu Fan, Zhenhua Guo 0003, Yaqian Zhao, RenGang Li |
IEEE Trans. Multim. | 7 |
| 2024 | A Survey on Performance Modeling and Prediction for Distributed DNN TrainingabstractThe recent breakthroughs in large-scale DNN attract significant attention from both academia and industry toward distributed DNN training techniques. Due to the time-consuming and expensive execution process of large-scale distributed DNN training, it is crucial to model and predict the performance of distributed DNN training before its actual deployment, in order to optimize the design of distributed DNN training at low cost. This paper analyzes and emphasizes the importance of modeling and predicting the performance of distributed DNN training, categorizes and analyses the related state-of-the-art works, and discusses future challenges and opportunities for this research field. The objectives of this paper are twofold: first, to assist researchers in understanding and choosing suitable modeling and prediction tools for large-scale distributed DNN training, and second, to encourage researchers to propose more valuable research about performance modeling and prediction for distributed DNN training in the future. Zhenhua Guo 0003, Yinan Tang, Jidong Zhai, Tongtong Yuan, Li Wang 0040, Yaqian Zhao, RenGang Li |
IEEE Trans. Parallel Distributed Syst. | 8 |
| 2023 | End-to-End Urban Autonomous Navigation with Decision Hindsight
Guangqing Liu, Ruyang Li, Qifu Hu, Yaqian Zhao, RenGang Li |
ICONIP (15) | 6 |
| 2023 | Enhancing Network by Reinforcement Learning and Neural Confined Local SearchabstractIt has been found that many real networks, such as power grids and the Internet, are non-robust, i.e., attacking a small set of nodes would cause the paralysis of the entire network. Thus, the Network Enhancement Problem~(NEP), i.e., improving the robustness of a given network by modifying its structure, has attracted increasing attention. Heuristics have been proposed to address NEP. However, a hand-engineered heuristic often has significant performance limitations. A recently proposed model solving NEP by reinforcement learning has shown superior performance than heuristics on in-distribution datasets. However, their model shows considerably inferior out-of-distribution generalization ability when enhancing networks against the degree-based targeted attack. In this paper, we propose a more effective model with stronger generalization ability by incorporating domain knowledge including measurements of local network structures and decision criteria of heuristics. We further design a hierarchical attention model to utilize the network structure directly, where the query range changes from local to global. Finally, we propose neural confined local search~(NCLS) to realize the effective search of a large neighborhood, which exploits a learned model to confine the neighborhood to avoid exhaustive enumeration. We conduct extensive experiments on synthetic and real networks to verify the ability of our models. Qifu Hu, Ruyang Li, Yaqian Zhao, RenGang Li |
IJCAI | 5 |
| 2023 | Context - Enhanced Meta-Reinforcement Learning with Data-Reused Adaptation for Urban Autonomous DrivingabstractAutonomous driving (AD) has experienced rapid development in recent years, and the reinforcement learning (RL) pipeline in trial-and-error manner can surpass human driving ability. However, the poor performance in sample efficiency and generalization limits RL applying in the challenging urban traffic scenarios. In this paper, we build a context-enhanced meta-RL framework with data-reused adaptation for challenging urban AD. At both the meta-learning and adaptation stages, the context-enhanced state representation is designed to reduce the perceptual gap in variant urban scenarios, improving the sample efficiency and robustness. At adaptation stage, the meta-training data with context-enhanced features are reused through propensity estimation to constrain the optimization objective of new tasks, aiming to maintain the good driving performance of meta-trained policy and fast adapt to the new tasks. Extensive experiments are conducted in CARLA simulator with various urban environments and task settings. The learning curves and quantitative comparisons validate the good sample efficiency and generalization of our proposed method, with state-of-the-art driving performance on urban AD benchmarks. Yaqian Zhao, RenGang Li, Qifu Hu, Tiejun Liu, Ruyang Li |
IJCNN | 3 |
| 2023 | Coresets based asynchronous network slimmingabstractAbstract Pruning is effective to reduce neural networks’ parameters and accelerate inferences, facilitating deep learning in resource-limited scenarios. This paper proposes an asynchronous pruning method for multi-branch networks on the basis of our previous work on channel coresets constructions, to achieve module-level pruning. Firstly, this paper accelerates coreset based pruning by batch sampling with a sampling probability decided on our-designed importance function. Secondly, this paper gives asynchronous pruning solutions with an in-place distillation of feature maps for deployment on multi-branch networks such as ResNet and SqueezeNet. Thirdly, this paper provides an extension to neuron pruning by grouping weights as channels. During tests on sensitivity of different layers to channel pruning, our method outperforms comparison schemes on object detection networks, indicating advantages of data-independent channel selections in maintaining precision. As shown in tests of asynchronous pruning solutions on multi-branch classification networks, our method further decreases FLOPs with a small accuracy decline on ResNet and acquires a small accuracy increment on SqueezeNet. In tests on neuron pruning, our method achieves an accuracy comparable to existing coreset based pruning methods by two solutions of precision recovery. Wenfeng Yin, Gang Dong, Yaqian Zhao, RenGang Li |
Appl. Intell. | 4 |
| 2023 | Correction to: Coresets based asynchronous network slimming
Wenfeng Yin, Gang Dong, Yaqian Zhao, RenGang Li |
Appl. Intell. | 4 |
| 2023 | Multi-agent deep reinforcement learning for online request scheduling in edge cooperation networks
Yaqiang Zhang, Ruyang Li, Yaqian Zhao, RenGang Li, Zhangbing Zhou |
Future Gener. Comput. Syst. | 4 |
| 2023 | Hierarchically stacked graph convolution for emotion recognition in conversationabstractAccurate emotion recognition can drive the robot to understand human affection intentions precisely and deliver the emotional response when communicating with a person. Recently, graph structure has been applied to explicitly capture the self and inter-dependencies of speakers in the conversation. However, the performance of the method is limited by inadequate discriminative information extraction based on naive graph convolution. In this paper, we propose a novel Hierarchically Stacked Graph Convolution Framework (HSGCF), which leverages hierarchical structure to extract emotional discriminative features. The proposed HSGCF uses five graph convolution layers connected hierarchically to establish a more discriminative emotional feature extractor. More importantly, to mitigate the over-smooth problem caused by deeper networks, Transformer structures with residual connection are introduced into HSGCF. Experimental results on the IEMOCAP benchmark dataset indicate the proposed framework achieves a 4.12% improvement in accuracy and a 4.80% improvement in F1 score compared with the baseline method. Binqiang Wang, Gang Dong, Yaqian Zhao, RenGang Li, Qichun Cao, Ke-Kun Hu, Dongdong Jiang |
Knowl. Based Syst. | 4 |
| 2022 | Context-Based Point Generation Network for Point Cloud Completion
Ruyang Li, Hui Wei 0005, Yaqian Zhao, RenGang Li |
ICONIP (1) | 5 |
| 2022 | Point Cloud Completion with Difference-Aware Point Voting
Ruyang Li, Hui Wei 0005, Yaqian Zhao, RenGang Li, Binqiang Wang |
ICONIP (6) | 5 |
| 2022 | Learning from Fourier: Leveraging Frequency Transformation for Emotion Recognition
Binqiang Wang, Gang Dong, Yaqian Zhao, RenGang Li |
ICONIP (2) | 4 |
| 2022 | Deep Reinforcement Learning based Mobility-Aware Service Migration for Multi-access Edge Computing EnvironmentabstractMulti-access Edge Computing (MEC) plays an im-portant role for providing end users with high reliability and low latency services at the edge of mobile network. In the scenario of Internet of Vehicles (IoV), vehicle users continually access nearby base stations to offload real-time tasks for reducing their computing overhead, while the ongoing services on current deployed edge nodes may be far away from users with the vehicles moving, potentially resulting in a high delay of data transmission. To address this challenge, in this paper, we propose a Deep Reinforcement Learning (DRL)-based mobility-aware service migration mechanism for effectively reducing the service delay and migration delay of the network. The proposed technique is adopted by re-calibrating required services at edge locations near the mobile user. Edge network state and user movement information are considered to ensure the generation of real-time service migration decision. Extensive experiments are conducted, and evaluation results demonstrate that our proposed DRL-based technique can effectively reduce the long-term average delay of the MEC system, compared with the state-of-the-art techniques. Yaqiang Zhang, RenGang Li, Yaqian Zhao, Ruyang Li |
ISCC | 2 |
| 2022 | Online Decentralized Task Allocation Optimization for Edge Collaborative NetworksabstractIn centralized task allocation strategies, real-time status information needs to be collected from distributed edge nodes. Therefore, the overloaded transmission on backbone network appears and leads to devastating decrease in the per-formance of centralized strategies. To address this issue, this paper proposes a multi-agent deep reinforcement learning based online decentralized task allocation mechanism, where each edge node makes task allocation decisions based on local network-state information. A centralized-training distributed-execution method is adopted to decrease data transmission load, and a value decomposition-based technique is applied at training stage for improving long-term performance of task allocation in edge col-laborative networks. Extensive experiments are conducted, and evaluation results demonstrate that our mechanism outperforms other three baseline algorithms in reducing the long-term average system delay and improving request completion rate. Yaqiang Zhang, Ruyang Li, Yaqian Zhao, RenGang Li, Xuelei Li |
ISCC | 4 |
| 2022 | Towards Further Comprehension on Referring Expression with RationaleabstractReferring Expression Comprehension (REC) is one important research branch in visual grounding, where the goal of REC is to localize a relevant object in the image, given an expression in the form of text to exactly describe a specific object. However, existing REC tasks aim at text content filtering and image object locating, which are evaluated based on the precision of the detection boxes. This may lead models to skip the learning process of multimodal comprehension directly and achieve good performance. In this paper, we work on how to enable an artificial agent to understand RE further and propose a more comprehensive task, called Further Comprehension on Referring Expression (FREC). In this task, we mainly focus on three sub-tasks: 1) correcting the erroneous text expression based on visual information; 2) generating the rationale of this input expression; 3) localizing the proper object based on the corrected expression. Accordingly, we make a new dataset named Further-RefCOCOs based on the RefCOCO, RefCOCO+, RefCOCOg benchmark datasets for this new task and make it publicly available. After that, we design a novel end-to-end pipeline to achieve these sub-tasks simultaneously. The experimental results demonstrate the validity of the proposed pipeline. We believe this work will motivate more researchers to explore along with this direction, and promote the development of visual grounding. RenGang Li, Baoyu Fan, Xiaochuan Li 0001, Zhenhua Guo 0003, Yaqian Zhao, Weifeng Gong, Endong Wang |
ACM Multimedia | 1 |
| 2022 | AI-VQA: Visual Question Answering based on Agent Interaction with InterpretabilityabstractVisual Question Answering (VQA) serves as a proxy for evaluating the scene understanding of an intelligent agent by answering questions about images. Most VQA benchmarks to date are focused on those questions that can be answered through understanding visual content in the scene, such as simple counting, visual attributes, and even a little challenging questions that require extra encyclopedic knowledge. However, humans have a remarkable capacity to reason dynamic interaction on the scene, which is beyond the literal content of an image and has not been investigated so far. In this paper, we propose Agent Interaction Visual Question Answering (AI-VQA), a task investigating deep scene understanding if the agent takes a certain action. For this task, a model not only needs to answer action-related questions but also to locate the objects in which the interaction occurs for guaranteeing it truly comprehends the action. Accordingly, we make a new dataset based on Visual Genome and ATOMIC knowledge graph, including more than 19,000 manually annotated questions, and will make it publicly available. Besides, we also provide an annotation of the reasoning path while developing the answer for each question. Based on the dataset, we further propose a novel method, called ARE, that can comprehend the interaction and explain the reason based on a given event knowledge base. Experimental results show that our proposed method outperforms the baseline by a clear margin. RenGang Li, Cong Xu 0001, Zhenhua Guo 0003, Baoyu Fan, Yaqian Zhao, Weifeng Gong, Endong Wang |
ACM Multimedia | 1 |
| 2022 | Non-Uniform Attention Network for Multi-modal Sentiment Analysis
Binqiang Wang, Gang Dong, Yaqian Zhao, RenGang Li, Qichun Cao, Yinyin Chao |
MMM (1) | 4 |
| 2021 | A fast, scalable, and energy-efficient edge acceleration architecture based on FPGA clusterabstractFPGA-based acceleration has been emerged to avoid the cloud computing overload problem by accelerating the compute-intensive workload on edge networks. Though existing studies for FPGA-based edge acceleration have focused on optimizing the computing time, they did not address the burden of the FPGA-aided server under an enormous computing requests circumstance. The massive computing requests can cause delays in transmission and processing at the FPGA-aided server, leading to long response times and high system energy consumption. Therefore, we propose an emerging edge acceleration architecture based on FPGA cluster over the low-latency RDMA-based network. Preliminary simulation results demonstrate that our architecture is fast, scalable, and energy-efficient in comparison with the FPGA-aided servers cluster. RenGang Li, Dongdong Su, Hongwei Kan |
CoNEXT | 1 |
| 2021 | You Get What You Sow: High Fidelity Image Synthesis with a Single Pretrained NetworkabstractState-of-the-art image synthesis methods are mostly based on generative adversarial networks and require large dataset and extensive training. Although the model-inversion-oriented branch of methods eliminate the training requirement, the quality of the resulting image tends to be limited due to the lack of sufficient natural and class-specific information. In this paper, we introduce a novel strategy for high fidelity image synthesis with a single pretrained classification network. The strategy includes a class-conditional natural regularization design and a corresponding metadata collecting procedure for different scenarios. We show that our method can synthesize high quality natural images that closely follow the features of one or more given seed images. Moreover, our method achieves surprisingly decent results in the task of sketch-based image synthesis without training. Finally, our method further improves the performance in terms of accuracy and efficiency in the data-free knowledge distillation task. Kefeng Zhu, Peilin Tong, Hongwei Kan, RenGang Li |
IJCAI | 4 |
| 2021 | Coresets Application in Channel Pruning for Fast Neural Network SlimmingabstractPruning reduces neural networks' parameters and accelerates inferences, enabling deep learning in resource-limited scenarios. Existing saliency-based pruning methods apply characteristics of feature maps or weights to judge the importance of neurons or structures, where weights' characteristics based methods are data-independent and robust for future input data. This paper proposes a coreset based pruning method for the data-independent structured compression, aiming to improve the construction efficiency of pruning. The first step of our method is to prune channels, according to the channel coreset merged from multi-rounds coresets constructions. Our method adjusts the importance function utilized in the random probability sampling during coresets construction procedures to achieve data-independent channel selections. The second step is recovering the precision of compressed networks through solving the compressed weights reconstruction by linear least squares. Our method is also generalized to implementations on multi-branch networks such as SqueezeNet and MobileNet-v2. In tests on classification networks like ResNet, it is observed that our method performs fast and achieves an accuracy decline as small as 0.99% when multiple layers are pruned without finetuning. As shown in evaluations on object detection networks, our method acquires the least decline in mAP indicator compared to comparison schemes, due to the advantage of data-independent channel selections of our method in preserving precision. Wenfeng Yin, Gang Dong, Yaqian Zhao, RenGang Li |
IJCNN | 4 |
| 2021 | Knowledge-Supervised Learning: Knowledge Consensus Constraints for Person Re-IdentificationabstractThe consensus of multiple views on the same data will provide extra regularization, thereby improving accuracy. Based on this idea, we proposed a novel Knowledge-Supervised Learning (KSL) method for person re-identification (Re-ID), which can improve the performance without introducing extra inference cost. Firstly, we introduce isomorphic auxiliary training strategy to conduct basic multiple views that simultaneously train multiple classifier heads of the same network on the same training data. The consensus constraints aim to maximize the agreement among multiple views. To introduce this regular constraint, inspired by knowledge distillation that paired branches can be trained collaboratively through mutual imitation learning. Three novel constraints losses are proposed to distill the knowledge that needs to be transferred across different branches: similarity of predicted classification probability for cosine space constraints, distance of embedding features for euclidean space constraints, hard sample mutual mining for hard sample space constraints. From different perspectives, these losses complement each other. Experiments on four mainstream Re-ID datasets show that a standard model with KSL method trained from scratch outperforms its ImageNet pre-training results by a clear margin. With KSL method, a lightweight model without ImageNet pre-training outperforms most large models. We expect that these discoveries can attract some attention from the current de facto paradigm of "pre-training and fine-tuning" in Re-ID task to the knowledge discovery during model training. Li Wang 0040, Baoyu Fan, Zhenhua Guo 0003, Yaqian Zhao, RenGang Li, Weifeng Gong, Endong Wang |
ACM Multimedia | 6 |
| 2021 | Deep Reinforcement Learning for DAG-based Concurrent Requests Scheduling in Edge Networks
Yaqiang Zhang, Ruyang Li, Zhangbing Zhou, Yaqian Zhao, RenGang Li |
WASA (3) | 5 |
| 2021 | An improved model training method for residual convolutional neural networks in deep learning
Xuelei Li, RenGang Li, Yaqian Zhao |
Multim. Tools Appl. | 2 |
| 2020 | Dense-Scale Feature Learning in Person Re-identification
Li Wang 0040, Baoyu Fan, Zhenhua Guo 0003, Yaqian Zhao, RenGang Li, Weifeng Gong |
ACCV (6) | 6 |
| 2020 | Contextual Multi-Scale Feature Learning for Person Re-IdentificationabstractRepresenting features at multiple scales is significant for person re-identification (Re-ID). Most existing methods learn the multi-scale features by stacking streams and convolutions without considering the cooperation of multiple scales at a granular level. However, most scales are more discriminative only when they integrate other scales as contextual information. We termed that contextual multi-scale. In this paper, we proposed a novel architecture, namely contextual multi-scale network (CMSNet), for learning common and contextual multi-scale representations simultaneously. The building block of CMSNet obtains contextual multi-scale representations by bidirectionally hierarchical connection groups: the forward hierarchical connection group for stepwise inter-scale information fusion and the backward hierarchical connection group for leap-frogging inter-scale information fusion. Too rich scale features without a selection will confuse the discrimination. Additionally, we introduced a new channel-wise scale selection module to dynamically select scale features for corresponding input image. To the best of our knowledge, CMSNet is the most lightweight model for person Re-ID and it achieves state-of-the-art performance on four commonly used Re-ID datasets, surpassing most large-scale models. Baoyu Fan, Li Wang 0040, Zhenhua Guo 0003, Yaqian Zhao, RenGang Li, Weifeng Gong |
ACM Multimedia | 6 |