Yaqian Zhao

dblp:171/5841 · DBLP profile ↗
← Back
53ranked-venue papers
0as first author
50since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 21 · 19 since 2021Artificial intelligence and machine learning · 18 · 17 since 2021Systems, architecture and hardware · 7 · 7 since 2021Computer networks · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Theory of computation · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Enabling Memory-Disaggregated Cloud Infrastructure for LLMs: An Adaptive CXL-based KV Cache Scheduling Approach
Yaqian Zhao, Yaqiang Zhang, Guangyuan Xu
INFOCOM2
2026 FPGA-based heterogeneous computing framework for environment-level parallel decision-making in autonomous driving and robotic control
Hongbin Yang 0003, Yaqian Zhao, Ruyang Li, Gang Dong
Expert Syst. Appl.3
2026 Enabling Learning-Based Efficiency Optimizer With Shadow Cycles in Resource-Constrained Autonomous Embedded Systems
abstract
The emerging trend of autonomous embedded systems (AES) is promising to minimize human intervention in critical tasks. In the pursuit of maximal per-watt performance, the complex hardware and software of AES require intelligent energy efficiency optimizers (EO), and the stochastic runtime variances require continuous EO. However, deploying the desirable ondevice EO causes severe performance slowdown due to contention on limited computing power with the AES pipeline. We find that there are ignored and underutilized heterogeneous resources within AES for costly EO, which results from unbalanced accelerator behaviors and misaligned parallel inference executions. We experimentally and theoretically analyze theShadow Cycleswithin the realistic autonomous Bird’s Eye View pipeline on commercial embedded platforms, categorizing them into vertical and horizontal types with distinct properties.In this paper, we introduceSHEEO+, a continuous and intelligent energy efficiency optimizer that utilizes ignored heterogeneous shadow cycles. It achieves continuous and lightweight AES monitoring with the observation module, as well as intelligent and efficient AES power management with the optimization module. On the one hand,SHEEO+ observes both the internal runtime status and external environment variance with portable interfaces to capture shadow cycles and real-time states. On the other hand,SHEEO+ optimizes power configurations per iteration based on deep reinforcement learning (DRL) methods. It tailors DRL for two types of shadow cycles and invocates optimization processes based on resource availability. To extensively evaluateSHEEO+, we implement a prototype and deploy it on realistic edge platforms. The evaluation results show thatSHEEO+ utilizes up to 74.2% shadow cycles and achieves up to 18.6% energy efficiency improvements compared to state-of-the-art energy efficiency optimizers with negligible deployment overheads.
Xinkai Wang 0003, Chao Li 0009, Xiaofeng Hou, Jing Wang 0055, Minyi Guo, Yaqian Zhao
IEEE Trans. Computers9
2025 MMEditor: Multimodal Prompt-Driven 3D Gaussian Splatting Editing
abstract
We propose a multimodal 3D scene editing framework MMEditor to create or modify objects within an extant 3D Gaussian Splatting (3DGS) according to text and image prompts. MMEditor employs a multimodal image editing module to iteratively optimize 3D Gaussians in editing regions for delicate and multi-view consistent 3D editing. The key multimodal image editing module can perform editing with accurate appearance and location control, which is achieved by two designs. First, a multimodel adapter block takes the reference image as a foreign language to augment the text prompt, enabling editing results to align with the generic text description and the unique characteristics in the reference image. Second, an attention-based localization block localizes cross-attention with user-defined 3D bounding boxes, thereby ensuring the editing occurs in editing regions. Experiments demonstrate that our method achieves more accurate and controllable results than previous state-of-the-art methods.
RenGang Li, Yaqian Zhao, Xiaohui Zhang 0017, Hui Wei 0005, Ruyang Li
ICASSP3
2025 Improving Height Prediction for Vision-Based Roadside 3D Object Detection
abstract
Roadside vision-based 3D object detection is vital in many applications, such as autonomous driving. The mainstream methods enhance the accuracy of distance estimation by converting predicted height distribution into depth distribution. However, predicting object’s height in roadside perception is challenging, particularly for distant and small objects. Therefore, this work proposes a series of methods to optimize height prediction. Firstly, we propose depth and height decoding supervision methods to optimize the height network by supervising the decoded depth and height values. Then, a height distribution alignment loss is introduced to optimize the height network by constraining the consistency of height distributions among objects of the same category. Experimental results demonstrate that the proposed optimization methods can effectively improve the accuracy of roadside vision-based 3D object detection. For instance, our methods improve the accuracies by 2.31%, 4.33%, and 4.24% for the cyclist category at easy, medium, and hard levels, respectively.
Tengfei Zhang 0004, RenGang Li, Yaqian Zhao, Ruyang Li
ICASSP4
2025 Dropletvideo: A Dataset and Approach to Explore Integral Spatio-Temporal Consistent Video Generation
Guoguang Du 0001, Xiaochuan Li 0001, Qi Jia 0004, Lu Liu 0009, Cong Xu 0001, Zhenhua Guo 0003, Yaqian Zhao, Xiaoli Gong, RenGang Li, Baoyu Fan
ICCV10
2025 Proactive Fault-tolerance Driven Task Scheduling System for IoV Edge Networks
abstract
The emergence of Internet of Vehicles (IoV) technology provides a wider range of application scenarios for edge computing based on Vehicle-to-everything (V2X). It is essential to ensure the high availability and reliability of services in IoV systems. Currently, cloud service providers have established a data center level of fault tolerance, such as redundancy and checkpoints, guaranteeing the reliability of cloud infrastructure and reducing phenomena such as service termination or downtime. However, current computing systems reactively handle failures. Especially in edge computing, this approach not only lacks flexibility but also consumes excessive system resources, which is not conducive to ensuring the reliability in resource-constrained systems and poses security risks to end users. To mitigate this problem, we propose a Proactive Fault-tolerance Driven Task Scheduling System. Different from the traditional reactive strategies, the proposed framework predicts the possible system crashes by monitoring the critical state indicators of the computing system. According to the prediction results, a class of tasks or services that are most likely to be terminated are rescheduled in advance. Extensive experiments are conducted, and evaluation results demonstrate that our proposed proactive fault tolerance framework can effectively improve the long-term performance of the IoV edge system.
Yaqiang Zhang, RenGang Li, Yaqian Zhao, Hongzhi Shi, Guangyuan Xu
ICNP3
2025 Repurpose Accel-Sim for Next Generation NVIDIA Jetson GPU Architectural Design
abstract
The growing adoption of NVIDIA Jetson devices in edge-AI applications highlights the need for accurate architecture simulation tools on their integrated GPUs. Existing cycle-accurate GPU simulators primarily target traditional discrete GPUs and exhibit significant inaccuracies when applied to Jetson integrated GPUs. While Accel-Sim serves as the most widely used academic simulator for NVIDIA GPU research, its lack of support for the latest Jetson integrated GPUs severely hinders architectural exploration for next generation edge-AI devices.We propose Accel-Sim-J, which bridges the gap by repurposing Accel-Sim simulation framework to NVIDIA Jetson GPUs. We refine three major Accel-Sim framework components by applying tuner modifications, GPGPU-Sim performance model enhancements, and correlator adjustments. These improvements enable precise Jetson GPU simulation support, reducing simulation cycle errors from 29.0% to 22.7% on the Rodinia benchmark and from 26.1% to 16.1% on a transformer block. Furthermore, our enhanced architectural support for Ampere GPUs achieves a considerable reduction in simulation error (from 140.1% to 50.2%) for GEMM kernels.Based on Accel-Sim-J, we conduct a case study investigating the architecture design difference between an edge GPU and a traditional one. Specifically, we compare the optimal Compute-to-Cache (C2C) ratio by changing the L2 cache size of Jetson AGX Orin and RTX 3090. We conclude that Jetson GPUs demonstrate a higher optimal C2C ratio than discrete GPUs for the same workloads. We suggest that designers reduce the on-chip area proportion of the L2 cache in the next generation Jetson GPU design for better performance and efficiency.
Chao Li 0009, Xiaofeng Hou, Yaqian Zhao, Jingwen Leng, Li Li 0012, Minyi Guo
ISLPED5
2025 Asymptotically Optimal Repair of Reed-Solomon Codes with Small Sub-Packetization under Rack-Aware Model
abstract
This paper presents a comprehensive study on the asymptotically optimal repair of Reed-Solomon (RS) codes with small sub-packetization, specifically tailored for rack-aware distributed storage systems. Through the utilization of multibase expansion, we introduce a novel approach that leverages monomials to construct linear repair schemes for RS codes. Our repair schemes which adapt to all admissible parameters achieve asymptotically optimal repair bandwidth while significantly reducing the sub-packetization compared with existing schemes. Furthermore, our approach is capable of repairing RS codes with asymptotically optimal repair bandwidth under the homogeneous storage model, achieving smaller sub-packetization than existing methods.
Zhongyan Liu, RenGang Li, Yaqian Zhao, Yaqiang Zhang
ITW4
2025 Tripartite interaction representation learning for multi-modal sentiment analysis
Binqiang Wang, Gang Dong, Yaqian Zhao, RenGang Li, Wenfeng Yin
Expert Syst. Appl.3
2025 A Multi-Granularity Relation Graph Aggregation Framework With Multimodal Clues for Social Relation Reasoning
abstract
The social relation is a fundamental attribute of human beings in daily life. The ability of humans to form large organizations and institutions stems directly from our complex social networks. Therefore, understanding social relationships in the context of multimedia is crucial for building domain-specific or general artificial intelligence systems. The key to reason social relations lies in understanding the human interactions between individuals through multimodal representations such as action and utterance. However, due to video editing techniques and various narrative sequences in videos, two individuals with social relationships may not appear together in the same frame or clip. Additionally, social relations may manifest in different levels of granularity in video expressions. Previous research has not effectively addressed these challenges. Therefore, this paper proposes aMulti-Granularity Relation Graph Aggregation Framework(MGRG) to enhance the inference ability for social relation reasoning in multimedia content, like video. Different from existing methods, our method considers the paradigm of jointly inferring the relations by constructing a social relation graph. We design a hierarchical multimodal relation graph illustrating the exchange of information between individuals' roles, capturing the complex interactions at multi-levels of granularity from fine to coarse. In MGRG, we propose two aggregation modules to cluster multimodal features in different granularity layer relation graph, considering temporal aspects and importance. Experimental results show that our method generates a logical and coherent social relation graph and improves the performance in accuracy.
Cong Xu 0001, Feiyu Chen 0005, Qi Jia 0004, Yunji Li, Yaqian Zhao, Changming Zhao
IEEE Trans. Multim.7
2025 Leveraging Graph Analysis to Pinpoint Root Causes of Scalability Issues for Parallel Applications
abstract
It is challenging to scale parallel applications to modern supercomputers because of load imbalance, resource contention, and communications between processes. Profiling and tracing are two main performance analysis approaches for detecting these scalability bottlenecks. Profiling is low-cost but lacks detailed dependence for identifying root causes. Tracing records plentiful information but incurs significant overheads. To address these issues, we presentScalAna, which employs static analysis techniques to combine the benefits of profiling and tracing - it enables tracing's analyzability with overhead similar to profiling.ScalAnauses static analysis to capture program structures and data dependence of parallel applications, and leverages lightweight profiling approaches to record performance data during runtime. Then a parallel performance graph is generated with both static and dynamic data. Based on this graph, we design a backtracking detection approach to automatically pinpoint the root causes of scaling issues. We evaluate the efficacy and efficiency ofScalAnausing several real applications with up to 704K lines of code and demonstrate that our approach can effectively pinpoint the root causes of scaling loss with an average overhead of 5.65% for up to 16,384 processes. By fixing the root causes detected by our tool, it achieves up to 33.01% performance improvement.
Yuyang Jin 0001, Haojie Wang 0004, Xiongchao Tang, Zhenhua Guo 0003, Yaqian Zhao, Torsten Hoefler, Tao Liu 0029, Xu Liu 0001, Jidong Zhai
IEEE Trans. Parallel Distributed Syst.5
2024 Image Content Generation with Causal Reasoning
abstract
The emergence of ChatGPT has once again sparked research in generative artificial intelligence (GAI). While people have been amazed by the generated results, they have also noticed the reasoning potential reflected in the generated textual content. However, this current ability for causal reasoning is primarily limited to the domain of language generation, such as in models like GPT-3. In visual modality, there is currently no equivalent research. Considering causal reasoning in visual content generation is significant. This is because visual information contains infinite granularity. Particularly, images can provide more intuitive and specific demonstrations for certain reasoning tasks, especially when compared to coarse-grained text. Hence, we propose a new image generation task called visual question answering with image (VQAI) and establish a dataset of the same name based on the classic Tom and Jerry animated series. Additionally, we develop a new paradigm for image generation to tackle the challenges of this task. Finally, we perform extensive experiments and analyses, including visualizations of the generated content and discussions on the potentials and limitations. The code and data are publicly available under the license of CC BY-NC-SA 4.0 for academic and non-commercial usage at: https://github.com/IEIT-AGI/MIX-Shannon/blob/main/projects/VQAI/lgd_vqai.md.
Xiaochuan Li 0001, Baoyu Fan, Zhenhua Guo 0003, Yaqian Zhao, RenGang Li
AAAI7
2024 A Distributed Algorithm for Rumor Blocking on Social Networks
Ruidong Yan, Zhenhua Guo 0003, Yaqian Zhao, RenGang Li, Xingjian Ding
COCOON (2)3
2024 DM-SARAH: A Variance Reduction Optimization Algorithm for Machine Learning Systems
abstract
Nowadays, the variance reduction (VR) technique is used to improve the performance of gradient-type algorithms in machine learning and deep learning. However, some existing VR algorithms require unrealistic assumptions or conditions such as τ-gradient dominated and Polyak-Lojasiewicz (PL) conditions, which limit their applications. In this paper, we present a Double Mini-batch StochAstic Recursive grAdient algoritHm (DM-SARAH) without these assumptions or conditions to solve the convex and non-convex optimization problems respectively. The main contributions of this paper are twofold: (1) At the theoretical level, we optimize the convergence rate and provide a complexity analysis of DM-SARAH, and (2) At the experimental level, we evaluate the effectiveness and efficiency of the proposed algorithm on various datasets. The experimental results indicate that the proposed algorithm outperforms existing methods.
RenGang Li, Ruidong Yan, Zhenhua Guo 0003, Zhi-Yong Qiu, Yaqian Zhao
GLOBECOM5
2024 Segment Anything Model Guided Semantic Knowledge Learning For Remote Sensing Change Detection
abstract
Existing deep learning based remote sensing change detection (RSCD) methods only rely on binary ground-truth to guide the network learning while neglecting the useful semantic guidance. As a result, the network can be readily misled by irrelevant category changes, leading to degraded performance and slow convergence of the model. To this end, we propose a novel segment anything model (SAM) guided framework, termed as SAM-CD, which mines the rich semantic knowledge from the SAM for RSCD. Specifically, we first employ a transformer encoder to extract multi-scale global features from the bi-temporal images. Meanwhile, we obtain semantic prior masks from the bi-temporal images by providing the SAM with category-relevant text prompts. Then, using the semantic prior masks as constraints, we design a masked attention module (MAM) that generates local features related to the interested categories. Finally, the local and global features are fused and fed into a multi-layer perception (MLP) decoder to obtain the change map. The whole network is trained in an end-to-end manner that can readily encode the rich semantic knowledge of the changed targets to predict an accurate change map. Extensive experiments demonstrate that the proposed SAM-CD achieves state-of-the-art performance on a variety of benchmark datasets.
Zixuan Sun, Huihui Song 0003, Kaihua Zhang 0001, Gang Dong, Lingyan Liang, Yaqian Zhao
ICASSP6
2024 Glance, Focus and Refinement Network for Remote Sensing Change Detection
abstract
Existing change detection (CD) methods often directly fuse the multi-level features from bi-temporal remote sensing images without discriminatively considering each pixel's importance. Despite the demonstrated success, unselectively mixing the features degrades the model's performance to effectively capture the change targets due to the imbalance ratio between the change regions and the whole scene. To this end, this paper presents a glance, focus, and refinement network (GFRNet), which formulates CD as a continuous, step-by-step focusing process to mimic the human visual system. Specifically, the GFRNet first employs a transformer encoder to extract the global features from the bi-temporal images, where each feature takes a glance at the whole scene. Then, the GFRNet gradually pays attention to a cascade of salient regions, and ultimately progressively refines its focus on the desired areas of change. Comprehensive evaluations on two extensively utilized benchmark datasets, including LEVIR-CD and WHU-CD, demonstrate the superiority of our GFR-Net to a variety of state-of-the-art methods.
Zixuan Sun, Yuhui Zheng, Kaihua Zhang 0001, Gang Dong, Lingyan Liang, Yaqian Zhao
ICASSP7
2024 FHNTT: a flexible Number Theoretic Transform design based on hybrid-radix butterfly
abstract
Emerging technologies, such as cloud computing and artificial intelligence, significantly arouse concern about data security and privacy. Homomorphic encryption (HE) is a promising invention, which enables computation on encrypted data without decrypting it so as to ensure data security and privacy. Nevertheless, computation within homomorphic encryption involves time-consuming operations, e.g., Number Theoretic Transform (NTT). The tremendous computation overhead is the critical obstacle in deploying HE applications widely. Besides, in order to meet the performance and security requirements of different applications, it is pivotal to design parametric NTT architecture. In this paper, we propose a flexible and parametric NTT accelerating scheme based on hybrid-radix butterfly, named FHNTT. Specifically, we construct high radix butterfly units and divide the computation of them into several stages such that every stage can be performed pipelined. The number of required twiddle factors declines with the increase of radix value. In addition, we adopt address offset strategy to reduce memory consumption. We implement FHNTT on FPGA due to its fine-grained parallel computing capabilities and customized architecture. Empirical results show that FHNTT has an improved performance compared with other NTT architectures and supports a wide range of parameters. Concretely, FHNTT achieves up to 1.99 × to 2.78 × improvement in latency over other FPGA implementations and the memory utilization rate is up to 94%. Moreover, the flexibility makes FHNTT applicable to multiple use cases.
RenGang Li, Yaqian Zhao, Ruyang Li, Zhiyuan Su, Xuelei Li
ISPA3
2024 A Pseudo-Hierarchical Planning Framework with Dynamic-Aware Reinforcement Learning for Autonomous Driving
abstract
Reinforcement Learning (RL) over motion skill space has been verified to generate more diverse behaviors than that over low-level control space, and has exhibited superior autonomous driving performance in complex traffic scenarios. However, the incomplete observations pose challenges in achieving efficient skill exploration under unsupervised conditions, hampering the driving performance and applicability. In this paper, we propose a dynamic-aware RL with hybrid network (Da-HnRL) to develop a pseudo-hierarchical planning framework for better motion skill learning in challenging dense traffics. Based on the semi-POMDP modeling, we construct a hybrid network with skip connections as the RL backbone, facilitating a better understanding of the underlying system dynamics. Then we design an efficiency-oriented reward shaping mechanism to incentivize active skill exploration, promoting enhanced trade-off between exploration and exploitation. Furthermore, we provide a comprehensive scoring mechanism for policy identification, ensuring the near-optimality. We validate the proposed methods on challenging dense-traffic tasks. The results demonstrate the superiority of our approach over previous methods, with improved learning efficiency, driving stability and generalization.
Yaqian Zhao, RenGang Li, Qifu Hu, Tengfei Zhang 0004, Ruyang Li
IV2
2024 Infer Induced Sentiment of Comment Response to Video: A New Task, Dataset and Baseline
abstract
Existing video multi-modal sentiment analysis mainly focuses on the sentiment expression of people within the video, yet often neglects the induced sentiment of viewers while watching the videos. Induced sentiment of viewers is essential for inferring the public response to videos and has broad application in analyzing public societal sentiment, effectiveness of advertising and other areas. The micro videos and the related comments provide a rich application scenario for viewers’ induced sentiment analysis. In light of this, we introduces a novel research task, Multimodal Sentiment Analysis for Comment Response of Video Induced(MSA-CRVI), aims to infer opinions and emotions according to comments response to micro video. Meanwhile, we manually annotate a dataset named Comment Sentiment toward to Micro Video (CSMV) to support this research. It is the largest video multi-modal sentiment dataset in terms of scale and video duration to our knowledge, containing 107, 267 comments and 8, 210 micro videos with a video duration of 68.83 hours. To infer the induced sentiment of comment should leverage the video content, we propose the Video Content-aware Comment Sentiment Analysis (VC-CSA) method as a baseline to address the challenges inherent in this new task. Extensive experiments demonstrate that our method is showing significant improvements over other established baselines. We make the dataset and source code publicly available at https://github.com/IEIT-AGI/MSA-CRVI.
Qi Jia 0004, Baoyu Fan, Cong Xu 0001, Lu Liu 0009, Guoguang Du 0001, Zhenhua Guo 0003, Yaqian Zhao, Xuanjing Huang 0001, RenGang Li
NeurIPS8
2024 Boosting Data Center Performance via Intelligently Managed Multi-backend Disaggregated Memory
abstract
Existing disaggregated memory (DM) systems face a problem of underutilized far memory bandwidth, which greatly limits the data throughput when processing data-intensive applications. Specifically, prior works all target runtime design for a single PCIe-based secondary memory device (i.e., single-backend far memory) with low data bandwidth and high system overhead. In this work, we take the first step to realize a well-crafted, multi-backend DM system with scale-out far memory paths. We propose xDM, a novel DM management scheme that can dynamically build and implicitly select appropriate far memory access paths. As part of xDM, we devise a smart far memory configuration strategy that can further optimize bandwidth usage effectiveness by tuning a wide set of key parameters based on synthesized information of application page data. Our design shows up to $3.9 \times$ data swap performance speedup, $2.8 \times$ data throughput increase, and $5.1 \times$ data center task throughput improvement compared with state-of-the-art works.
Jing Wang 0055, Hanzhang Yang, Chao Li 0009, Yiming Zhuansun, Wang Yuan, Xiaofeng Hou, Minyi Guo, Yang Hu 0001, Yaqian Zhao
SC10
2024 Group-wise co-salient object detection via multi-view self-labeling novel class discovery
Gang Dong, Lingyan Liang, Yaqian Zhao, Kaihua Zhang 0001
Frontiers Comput. Sci.4
2024 MVIndEmo: a dataset for micro video public-induced emotion prediction on social media
abstract
Abstract Distinct from the realm of perceived emotion research, induced emotion pertains to the emotional responses engendered within content consumers. This facet has garnered considerable attention and finds extensive application in the analysis of public social media. However, the advent of micro videos presents unique challenges when attempting to discern the induced emotional patterns exhibited by content consumers, owing to their free-style representation and other factors. Consequently, we have put forth two novel tasks concerning the recognition of public-induced emotion on micro videos: emotion polarity and emotion classification. Additionally, we have introduced a accessible dataset specifically tailored for the analysis of public-induced emotion on micro videos. The data corpus has been meticulously collected from Tiktok, a burgeoning social media platform renowned for its trendsetting content. To construct the dataset, we have selected eight captivating topics that elicit vibrant social discussions. In devising our label generation strategy, we have employed an automated approach characterized by the fusion of multiple expert models. This strategy incorporates a confidence measure method that relies on three distinct models for effectively aggregating user comments. To accommodate adaptable benchmark configurations, we provide both binary classification labels and probability distribution labels. The dataset encompasses a vast collection of 7,153 labeled micro videos. We have undertaken an extensive statistical analysis of the dataset to provide a comprehensive overview composition. It is our earnest aspiration that this dataset will serve as a catalyst for pioneering research avenues in the analysis of emotional patterns and the understanding of multi-modal information.
Zhenhua Guo 0003, Qi Jia 0004, Baoyu Fan, Cong Xu 0001, Yaqian Zhao, RenGang Li
Multim. Syst.7
2024 Easy Pruning via Coresets and Structural Re-Parameterization
abstract
Inference time pruning is characteristic in high construction efficiency, since it dramatically reduces the dependency on finetuning to recover precision. It is adequate to reconstruct compressed convolution kernels by optimizing the loss of feature map reconstruction. However, the accuracy decline of compressed network increases as the loss of feature map reconstruction accumulates layer by layer. To enhance layerwise convolution kernel reconstruction, this paper proposes a hybrid method via combining coresets theory and structural re-parameterization, enabling shallow transfer learning (STL) during inference time pruning. Firstly, our method achieves STL by implementing structural re-parameterization in the process of convolution kernel reconstruction, to adapt to effects of one layer's reconstruction loss on the next layers' inputs. Secondly, a channel-wise scaling process is designed on the basis of coresets theory, to enhance approximation in the mapping from drifted inputs to original feature maps. Selectively, a maximum mean discrepancy based decision-making process is built for switching in two patterns of our method. Tests are executed on image classification and arrhythmia detection. As observed on ImageNet datasets, coresets theory based scaling is more effective at filter level for DenseNet and MobileNet-v2 and resultful at unit kernel level for ResNet and SqueezeNet.
Wenfeng Yin, Gang Dong, Dianzheng An, Yaqian Zhao, Binqiang Wang
IEEE Signal Process. Lett.4
2024 Inexactly Matched Referring Expression Comprehension With Rationale
abstract
Referring Expression Comprehension (REC) is a multimodal comprehension task that aims to locate an object in an image, given a text description. Traditionally, during the existing REC tasks, there has been a basic assumption that the given text expression and the image are usually exactly matched to each other. However, in real-world scenarios, there is uncertainty in how well the image and text match each other exactly. Illegible objects in the image or ambiguous phrases in the text have the potential to significantly degrade the performance of conventional REC tasks. To overcome these limitations, we consider a more practical and comprehensive REC task, where the given image and its referring text expression can be inexactly matched. Our models aim to correct such inexact matching and supply corresponding interpretations. We refer to this task asFurther REC (FREC). This task is divided into three subtasks: 1) correcting the erroneous text expression using visual information, 2) generating the rationale for this input expression, and 3) localizing the proper object based on the corrected expression. We introduce three new datasets for FREC:Further-RefCOCOs,Further-CopsrefandFurther-Talk2Car. These datasets are based on the existing REC datasets, including RefCOCO and Talk2Car. We developed a novel pipeline architecture to execute the three subtasks simultaneously in an end-to-end fashion. Next, we developed an elastic masked language modeling (EMLM) training head to rectify text errors with uncertain lengths. Our experimental results demonstrate the validity of our proposed pipeline. We hope this work sparks more research focused on inexactly matched REC.
Xiaochuan Li 0001, Baoyu Fan, Zhenhua Guo 0003, Yaqian Zhao, RenGang Li
IEEE Trans. Multim.6
2024 A Survey on Performance Modeling and Prediction for Distributed DNN Training
abstract
The recent breakthroughs in large-scale DNN attract significant attention from both academia and industry toward distributed DNN training techniques. Due to the time-consuming and expensive execution process of large-scale distributed DNN training, it is crucial to model and predict the performance of distributed DNN training before its actual deployment, in order to optimize the design of distributed DNN training at low cost. This paper analyzes and emphasizes the importance of modeling and predicting the performance of distributed DNN training, categorizes and analyses the related state-of-the-art works, and discusses future challenges and opportunities for this research field. The objectives of this paper are twofold: first, to assist researchers in understanding and choosing suitable modeling and prediction tools for large-scale distributed DNN training, and second, to encourage researchers to propose more valuable research about performance modeling and prediction for distributed DNN training in the future.
Zhenhua Guo 0003, Yinan Tang, Jidong Zhai, Tongtong Yuan, Li Wang 0040, Yaqian Zhao, RenGang Li
IEEE Trans. Parallel Distributed Syst.7
2023 Group-Wise Co-Salient Object Detection with Siamese Transformers Via Brownian Distance Covariance Matching
abstract
Co-salient object detection (CoSOD) aims to discover and segment foreground targets in a group of images with the same semantic category. Existing mainstream approaches often employ convolutional neural networks (CNNs) to learn the semantic-invariant features from a group of images. Despite demonstrated success, there exist two limitations: 1) The CNNs introduce the inductive bias of locality that are difficult to model long-range dependency, limiting their feature representation capability. 2) Their models lack discriminability to differentiate semantic differences between different groups since only one group of images with the same semantic category has been taken into account for model training. To address these issues, this paper presents a Siamese Transformer architecture for CoSOD that can fully mine the group-wise semantic contrast information for more discriminative feature learning. Specifically, the designed Siamese Transformer takes two groups of images as input for feature contrastive learning. Each group is processed by a Transformer branch with shared weights to capture the long-range interaction information. Besides, to model the complex non-linear interactions between these two branches, we further design a Brownian distance covariance (BDC) module that uses joint distribution to measure the inter- and intra-group semantic similarity. The BDC can be efficiently calculated in closed form that can fully characterize independence for effective feature contrastive learning. Extensive evaluations on the three largest and most challenging benchmark datasets (CoSal2015, CoCA, and CoSOD3k) demonstrate the superiority of our method over a variety of state-of-the-art methods.
Lingyan Liang, Yaqian Zhao, Kaihua Zhang 0001
ICASSP4
2023 Object-Aware Calibrated Depth-Guided Transformer for RGB-D Co-Salient Object Detection
abstract
The key role of RGB-D co-salient object detection is to effectively fuse the common information of RGB and depth signals. Existing works directly mix the information captured from both original depth maps and RGB images, but ignore one critical issue: due to the low contrast of the neighborhood objects in depth, the depth maps’ salient regions may correspond to the interference background regions in the RGB images, thereby leading to unsatisfying performance. To address this issue, we propose an Object-aware Calibrated Depth guided transformer (dubbed as OCDFormer) for RGB-D co-salient object detection. The OCDFormer mainly consists of two key designs: First, we design a depth calibration module via spectral clustering, which yields a group of calibrated depth maps that can highlight the co-object region while suppressing the interference regions. Second, we construct a cross-modal transformer, in which the common information from the RGB and the calibrated depth maps are fully captured by first injecting common tokens into the individual tokens, and then mixing them with an interaction-attention mechanism. Extensive evaluations demonstrate that our OCDFormer sets a new state-of-the-art on two public standard benchmarks including RGB-D CoSall5O and RGB-D CoSegl83.
Lingyan Liang, Yaqian Zhao, Kaihua Zhang 0001
ICME3
2023 End-to-End Urban Autonomous Navigation with Decision Hindsight
Guangqing Liu, Ruyang Li, Qifu Hu, Yaqian Zhao, RenGang Li
ICONIP (15)5
2023 Enhancing Network by Reinforcement Learning and Neural Confined Local Search
abstract
It has been found that many real networks, such as power grids and the Internet, are non-robust, i.e., attacking a small set of nodes would cause the paralysis of the entire network. Thus, the Network Enhancement Problem~(NEP), i.e., improving the robustness of a given network by modifying its structure, has attracted increasing attention. Heuristics have been proposed to address NEP. However, a hand-engineered heuristic often has significant performance limitations. A recently proposed model solving NEP by reinforcement learning has shown superior performance than heuristics on in-distribution datasets. However, their model shows considerably inferior out-of-distribution generalization ability when enhancing networks against the degree-based targeted attack. In this paper, we propose a more effective model with stronger generalization ability by incorporating domain knowledge including measurements of local network structures and decision criteria of heuristics. We further design a hierarchical attention model to utilize the network structure directly, where the query range changes from local to global. Finally, we propose neural confined local search~(NCLS) to realize the effective search of a large neighborhood, which exploits a learned model to confine the neighborhood to avoid exhaustive enumeration. We conduct extensive experiments on synthetic and real networks to verify the ability of our models.
Qifu Hu, Ruyang Li, Yaqian Zhao, RenGang Li
IJCAI4
2023 Context - Enhanced Meta-Reinforcement Learning with Data-Reused Adaptation for Urban Autonomous Driving
abstract
Autonomous driving (AD) has experienced rapid development in recent years, and the reinforcement learning (RL) pipeline in trial-and-error manner can surpass human driving ability. However, the poor performance in sample efficiency and generalization limits RL applying in the challenging urban traffic scenarios. In this paper, we build a context-enhanced meta-RL framework with data-reused adaptation for challenging urban AD. At both the meta-learning and adaptation stages, the context-enhanced state representation is designed to reduce the perceptual gap in variant urban scenarios, improving the sample efficiency and robustness. At adaptation stage, the meta-training data with context-enhanced features are reused through propensity estimation to constrain the optimization objective of new tasks, aiming to maintain the good driving performance of meta-trained policy and fast adapt to the new tasks. Extensive experiments are conducted in CARLA simulator with various urban environments and task settings. The learning curves and quantitative comparisons validate the good sample efficiency and generalization of our proposed method, with state-of-the-art driving performance on urban AD benchmarks.
Yaqian Zhao, RenGang Li, Qifu Hu, Tiejun Liu, Ruyang Li
IJCNN2
2023 Aesthetics-Driven Virtual Time-Lapse Photography Generation
abstract
Time-lapse videos can visualize the temporal change of dynamic scenes and present wonderful sights with drastic variance in color appearance and rapid movement that interests people. We propose an aesthetics-driven virtual time-lapse photography framework to explore the automatic generation of time-lapse videos in the virtual world, which has potential applications like artistic creation and entertainment in the virtual space. We first define shooting parameters to parameterize the time-lapse photography process and accordingly propose image, video, and time-lapse aesthetic assessments to optimize these parameters, enabling the process to be autonomous and adaptive. We also build an interactive interface to visualize the shooting process and help users conduct virtual time-lapse photography by personalizing shooting parameters according to their aesthetic preferences. Finally, we present a two-stream time-lapse aesthetic model and a time-lapse aesthetic dataset, which can evaluate the aesthetic quality of time-lapse videos. Experimental results demonstrate our method can automatically generate time-lapse videos comparable to those of professional photographers and is more efficient.
Hui Wei 0005, Xin Jin 0015, Yihao Zhang 0010, Boyan Dong, Longteng Jiang, Xiaohui Zhang 0017, Ruyang Li, Yaqian Zhao
ACM Multimedia9
2023 Coresets based asynchronous network slimming
abstract
Abstract Pruning is effective to reduce neural networks’ parameters and accelerate inferences, facilitating deep learning in resource-limited scenarios. This paper proposes an asynchronous pruning method for multi-branch networks on the basis of our previous work on channel coresets constructions, to achieve module-level pruning. Firstly, this paper accelerates coreset based pruning by batch sampling with a sampling probability decided on our-designed importance function. Secondly, this paper gives asynchronous pruning solutions with an in-place distillation of feature maps for deployment on multi-branch networks such as ResNet and SqueezeNet. Thirdly, this paper provides an extension to neuron pruning by grouping weights as channels. During tests on sensitivity of different layers to channel pruning, our method outperforms comparison schemes on object detection networks, indicating advantages of data-independent channel selections in maintaining precision. As shown in tests of asynchronous pruning solutions on multi-branch classification networks, our method further decreases FLOPs with a small accuracy decline on ResNet and acquires a small accuracy increment on SqueezeNet. In tests on neuron pruning, our method achieves an accuracy comparable to existing coreset based pruning methods by two solutions of precision recovery.
Wenfeng Yin, Gang Dong, Yaqian Zhao, RenGang Li
Appl. Intell.3
2023 Correction to: Coresets based asynchronous network slimming
Wenfeng Yin, Gang Dong, Yaqian Zhao, RenGang Li
Appl. Intell.3
2023 Multi-agent deep reinforcement learning for online request scheduling in edge cooperation networks
Yaqiang Zhang, Ruyang Li, Yaqian Zhao, RenGang Li, Zhangbing Zhou
Future Gener. Comput. Syst.3
2023 Hierarchically stacked graph convolution for emotion recognition in conversation
abstract
Accurate emotion recognition can drive the robot to understand human affection intentions precisely and deliver the emotional response when communicating with a person. Recently, graph structure has been applied to explicitly capture the self and inter-dependencies of speakers in the conversation. However, the performance of the method is limited by inadequate discriminative information extraction based on naive graph convolution. In this paper, we propose a novel Hierarchically Stacked Graph Convolution Framework (HSGCF), which leverages hierarchical structure to extract emotional discriminative features. The proposed HSGCF uses five graph convolution layers connected hierarchically to establish a more discriminative emotional feature extractor. More importantly, to mitigate the over-smooth problem caused by deeper networks, Transformer structures with residual connection are introduced into HSGCF. Experimental results on the IEMOCAP benchmark dataset indicate the proposed framework achieves a 4.12% improvement in accuracy and a 4.80% improvement in F1 score compared with the baseline method.
Binqiang Wang, Gang Dong, Yaqian Zhao, RenGang Li, Qichun Cao, Ke-Kun Hu, Dongdong Jiang
Knowl. Based Syst.3
2023 Bi-RRNet: Bi-level recurrent refinement network for camouflaged object detection
Yan Liu 0004, Kaihua Zhang 0001, Yaqian Zhao, Qingshan Liu 0001
Pattern Recognit.3
2022 Context-Based Point Generation Network for Point Cloud Completion
Ruyang Li, Hui Wei 0005, Yaqian Zhao, RenGang Li
ICONIP (1)4
2022 Point Cloud Completion with Difference-Aware Point Voting
Ruyang Li, Hui Wei 0005, Yaqian Zhao, RenGang Li, Binqiang Wang
ICONIP (6)4
2022 Learning from Fourier: Leveraging Frequency Transformation for Emotion Recognition
Binqiang Wang, Gang Dong, Yaqian Zhao, RenGang Li
ICONIP (2)3
2022 Deep Reinforcement Learning based Mobility-Aware Service Migration for Multi-access Edge Computing Environment
abstract
Multi-access Edge Computing (MEC) plays an im-portant role for providing end users with high reliability and low latency services at the edge of mobile network. In the scenario of Internet of Vehicles (IoV), vehicle users continually access nearby base stations to offload real-time tasks for reducing their computing overhead, while the ongoing services on current deployed edge nodes may be far away from users with the vehicles moving, potentially resulting in a high delay of data transmission. To address this challenge, in this paper, we propose a Deep Reinforcement Learning (DRL)-based mobility-aware service migration mechanism for effectively reducing the service delay and migration delay of the network. The proposed technique is adopted by re-calibrating required services at edge locations near the mobile user. Edge network state and user movement information are considered to ensure the generation of real-time service migration decision. Extensive experiments are conducted, and evaluation results demonstrate that our proposed DRL-based technique can effectively reduce the long-term average delay of the MEC system, compared with the state-of-the-art techniques.
Yaqiang Zhang, RenGang Li, Yaqian Zhao, Ruyang Li
ISCC3
2022 Online Decentralized Task Allocation Optimization for Edge Collaborative Networks
abstract
In centralized task allocation strategies, real-time status information needs to be collected from distributed edge nodes. Therefore, the overloaded transmission on backbone network appears and leads to devastating decrease in the per-formance of centralized strategies. To address this issue, this paper proposes a multi-agent deep reinforcement learning based online decentralized task allocation mechanism, where each edge node makes task allocation decisions based on local network-state information. A centralized-training distributed-execution method is adopted to decrease data transmission load, and a value decomposition-based technique is applied at training stage for improving long-term performance of task allocation in edge col-laborative networks. Extensive experiments are conducted, and evaluation results demonstrate that our mechanism outperforms other three baseline algorithms in reducing the long-term average system delay and improving request completion rate.
Yaqiang Zhang, Ruyang Li, Yaqian Zhao, RenGang Li, Xuelei Li
ISCC3
2022 Towards Further Comprehension on Referring Expression with Rationale
abstract
Referring Expression Comprehension (REC) is one important research branch in visual grounding, where the goal of REC is to localize a relevant object in the image, given an expression in the form of text to exactly describe a specific object. However, existing REC tasks aim at text content filtering and image object locating, which are evaluated based on the precision of the detection boxes. This may lead models to skip the learning process of multimodal comprehension directly and achieve good performance. In this paper, we work on how to enable an artificial agent to understand RE further and propose a more comprehensive task, called Further Comprehension on Referring Expression (FREC). In this task, we mainly focus on three sub-tasks: 1) correcting the erroneous text expression based on visual information; 2) generating the rationale of this input expression; 3) localizing the proper object based on the corrected expression. Accordingly, we make a new dataset named Further-RefCOCOs based on the RefCOCO, RefCOCO+, RefCOCOg benchmark datasets for this new task and make it publicly available. After that, we design a novel end-to-end pipeline to achieve these sub-tasks simultaneously. The experimental results demonstrate the validity of the proposed pipeline. We believe this work will motivate more researchers to explore along with this direction, and promote the development of visual grounding.
RenGang Li, Baoyu Fan, Xiaochuan Li 0001, Zhenhua Guo 0003, Yaqian Zhao, Weifeng Gong, Endong Wang
ACM Multimedia7
2022 AI-VQA: Visual Question Answering based on Agent Interaction with Interpretability
abstract
Visual Question Answering (VQA) serves as a proxy for evaluating the scene understanding of an intelligent agent by answering questions about images. Most VQA benchmarks to date are focused on those questions that can be answered through understanding visual content in the scene, such as simple counting, visual attributes, and even a little challenging questions that require extra encyclopedic knowledge. However, humans have a remarkable capacity to reason dynamic interaction on the scene, which is beyond the literal content of an image and has not been investigated so far. In this paper, we propose Agent Interaction Visual Question Answering (AI-VQA), a task investigating deep scene understanding if the agent takes a certain action. For this task, a model not only needs to answer action-related questions but also to locate the objects in which the interaction occurs for guaranteeing it truly comprehends the action. Accordingly, we make a new dataset based on Visual Genome and ATOMIC knowledge graph, including more than 19,000 manually annotated questions, and will make it publicly available. Besides, we also provide an annotation of the reasoning path while developing the answer for each question. Based on the dataset, we further propose a novel method, called ARE, that can comprehend the interaction and explain the reason based on a given event knowledge base. Experimental results show that our proposed method outperforms the baseline by a clear margin.
RenGang Li, Cong Xu 0001, Zhenhua Guo 0003, Baoyu Fan, Yaqian Zhao, Weifeng Gong, Endong Wang
ACM Multimedia7
2022 Non-Uniform Attention Network for Multi-modal Sentiment Analysis
Binqiang Wang, Gang Dong, Yaqian Zhao, RenGang Li, Qichun Cao, Yinyin Chao
MMM (1)3
2021 Coresets Application in Channel Pruning for Fast Neural Network Slimming
abstract
Pruning reduces neural networks' parameters and accelerates inferences, enabling deep learning in resource-limited scenarios. Existing saliency-based pruning methods apply characteristics of feature maps or weights to judge the importance of neurons or structures, where weights' characteristics based methods are data-independent and robust for future input data. This paper proposes a coreset based pruning method for the data-independent structured compression, aiming to improve the construction efficiency of pruning. The first step of our method is to prune channels, according to the channel coreset merged from multi-rounds coresets constructions. Our method adjusts the importance function utilized in the random probability sampling during coresets construction procedures to achieve data-independent channel selections. The second step is recovering the precision of compressed networks through solving the compressed weights reconstruction by linear least squares. Our method is also generalized to implementations on multi-branch networks such as SqueezeNet and MobileNet-v2. In tests on classification networks like ResNet, it is observed that our method performs fast and achieves an accuracy decline as small as 0.99% when multiple layers are pruned without finetuning. As shown in evaluations on object detection networks, our method acquires the least decline in mAP indicator compared to comparison schemes, due to the advantage of data-independent channel selections of our method in preserving precision.
Wenfeng Yin, Gang Dong, Yaqian Zhao, RenGang Li
IJCNN3
2021 Knowledge-Supervised Learning: Knowledge Consensus Constraints for Person Re-Identification
abstract
The consensus of multiple views on the same data will provide extra regularization, thereby improving accuracy. Based on this idea, we proposed a novel Knowledge-Supervised Learning (KSL) method for person re-identification (Re-ID), which can improve the performance without introducing extra inference cost. Firstly, we introduce isomorphic auxiliary training strategy to conduct basic multiple views that simultaneously train multiple classifier heads of the same network on the same training data. The consensus constraints aim to maximize the agreement among multiple views. To introduce this regular constraint, inspired by knowledge distillation that paired branches can be trained collaboratively through mutual imitation learning. Three novel constraints losses are proposed to distill the knowledge that needs to be transferred across different branches: similarity of predicted classification probability for cosine space constraints, distance of embedding features for euclidean space constraints, hard sample mutual mining for hard sample space constraints. From different perspectives, these losses complement each other. Experiments on four mainstream Re-ID datasets show that a standard model with KSL method trained from scratch outperforms its ImageNet pre-training results by a clear margin. With KSL method, a lightweight model without ImageNet pre-training outperforms most large models. We expect that these discoveries can attract some attention from the current de facto paradigm of "pre-training and fine-tuning" in Re-ID task to the knowledge discovery during model training.
Li Wang 0040, Baoyu Fan, Zhenhua Guo 0003, Yaqian Zhao, RenGang Li, Weifeng Gong, Endong Wang
ACM Multimedia4
2021 Deep Reinforcement Learning for DAG-based Concurrent Requests Scheduling in Edge Networks
Yaqiang Zhang, Ruyang Li, Zhangbing Zhou, Yaqian Zhao, RenGang Li
WASA (3)4
2021 Classification of Remotely Sensed Images Using an Ensemble of Improved Convolutional Network
abstract
In the last few years, the deep learning methods, especially the residual neural network, have achieved impressive performance in remote sensing image recognition tasks. However, there are still specific problems that need to be addressed. It is well known that the first several layers of the network provide much discriminative information, and the ResNet reduces the size of the feature map so quickly that it failed to fully learn the information beneficial to classification in the early stage. Second, insufficient labeling data in remote sensing database may easily lead to overfitting and affect the final classification accuracy. Third, the optimal results cannot be achieved by relying solely on transfer learning. To overcome the problems mentioned earlier, we propose an enhanced residual neural network (ERNet) to improve the classification performance on remote sensing images. We moderately broadened the first several layers of the network, changed the size of the convolution filters, and made it learn more information of image features. Second, we add dropout layer to each residual unit of the proposed network to improve the accuracy and generalization power of ERNet. Finally, an ensemble of learning methods based on ERNet was introduced to improve the classification performance by fusing features of other baseline methods. Extensive experimental results on several benchmark data sets of remote sensing images demonstrate the superior performance of our proposed algorithm.
Li Wang 0040, Yanjiang Wang 0001, Yaqian Zhao, Baodi Liu
IEEE Geosci. Remote. Sens. Lett.3
2021 An improved model training method for residual convolutional neural networks in deep learning
Xuelei Li, RenGang Li, Yaqian Zhao
Multim. Tools Appl.3
2020 Dense-Scale Feature Learning in Person Re-identification
Li Wang 0040, Baoyu Fan, Zhenhua Guo 0003, Yaqian Zhao, RenGang Li, Weifeng Gong
ACCV (6)4
2020 Ensembled Tricks for Instance Segmentation
abstract
Computer Vision has attracted more and more attention with the fast development of deep learning. The instance segmentation area, which extends the Object detection, can better help us comprehend the surrounding environments. In this paper, we ensembled the tricks that can strengthen the model performance for instance segmentation. We do the ablation experiments for the MS-COCO datasets and LVIS datasets. The results demonstrate that the selected tricks can greatly boost the performance. With our tricks, our model achieves the 7th on the LVIS Challenge Track for ICCV 2019 workshop.
Yongfang Chen, Zhenhua Guo 0003, Yaqian Zhao
IWCMC6
2020 Contextual Multi-Scale Feature Learning for Person Re-Identification
abstract
Representing features at multiple scales is significant for person re-identification (Re-ID). Most existing methods learn the multi-scale features by stacking streams and convolutions without considering the cooperation of multiple scales at a granular level. However, most scales are more discriminative only when they integrate other scales as contextual information. We termed that contextual multi-scale. In this paper, we proposed a novel architecture, namely contextual multi-scale network (CMSNet), for learning common and contextual multi-scale representations simultaneously. The building block of CMSNet obtains contextual multi-scale representations by bidirectionally hierarchical connection groups: the forward hierarchical connection group for stepwise inter-scale information fusion and the backward hierarchical connection group for leap-frogging inter-scale information fusion. Too rich scale features without a selection will confuse the discrimination. Additionally, we introduced a new channel-wise scale selection module to dynamically select scale features for corresponding input image. To the best of our knowledge, CMSNet is the most lightweight model for person Re-ID and it achieves state-of-the-art performance on four commonly used Re-ID datasets, surpassing most large-scale models.
Baoyu Fan, Li Wang 0040, Zhenhua Guo 0003, Yaqian Zhao, RenGang Li, Weifeng Gong
ACM Multimedia5