Yongqiang Yao

dblp:226/6918 · DBLP profile ↗
← Back
18ranked-venue papers
4as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 3 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 1 · 1 first-author
YearPublicationVenuePosition
2026 Multi-objective parallel feasible direction algorithm for hypergraph partitioning problem with rank-two semidefinite programming relaxation
Yingying Li 0011, Yongqiang Yao, Hongwei Liu 0001
Integr.2
2025 TokenSim: Enabling Hardware and Software Exploration for Large Language Model Inference Systems
Feiyang Wu, Zhuohang Bian, Guoyang Duan, Tianle Xu, Junchi Wu, Yongqiang Yao, Ruihao Gong, Youwei Zhuo
APPT7
2025 Tool Playgrounds: A Comprehensive and Analyzable Benchmark for LLM Tool Invocation
abstract
The rapid advancement of large language models (LLMs) has paved the way for their use in solving real-world problems, which in turn has significantly driven the development of tool-assisted LLMs. This progress necessitates thorough evaluation methods. However, existing benchmarks typically only provide end-to-end scores but lack in-depth analysis and often suffer from issues such as instability. To address this gap, we have meticulously designed the Tool Playgrounds framework, a comprehensive, analyzable, and extensible benchmark. This framework evaluates boundary dimensions such as parameter missing interaction, parameter correction, tool failover, and leveraging internal knowledge. Our findings indicate that even the most advanced commercial models frequently overlook these essential aspects and face challenges in managing complex tool usage. To foster further research and development, we have made our code, dataset, and leaderboard publicly available on https://github.com/zhiwei-dong/ToolPlaygrounds.
Zhiwei Dong, Ruihao Gong, Yang Yong, Yongqiang Yao, Song-Lu Chen, Xu-Cheng Yin
ICASSP5
2025 OmniBal: Towards Fast Instruction-Tuning for Vision-Language Models via Omniverse Computation Balance
abstract
Vision-language instruction-tuning models have recently achieved significant performance improvements. In this work, we discover that large-scale 3D parallel training on those models leads to an imbalanced computation load across different devices. The vision and language parts are inherently heterogeneous: their data distribution and model architecture differ significantly, which affects distributed training efficiency. To address this issue, we rebalance the computational load from data, model, and memory perspectives, achieving more balanced computation across devices. Specifically, for the data, instances are grouped into new balanced mini-batches within and across devices. A search-based method is employed for the model to achieve a more balanced partitioning. For memory optimization, we adaptively adjust the re-computation strategy for each partition to utilize the available memory fully. These three perspectives are not independent but are closely connected, forming an omniverse balanced training framework. Extensive experiments are conducted to validate the effectiveness of our method. Compared with the open-source training code of InternVL-Chat, training time is reduced greatly, achieving about 1.8$\times$ speed-up. Our method’s efficacy and generalizability are further validated across various models and datasets. Codes will be released at https://github.com/ModelTC/OmniBal.
Yongqiang Yao, Jingru Tan, Feizhao Zhang, Yazhe Niu, Xin Jin 0008, Bo Li 0126, Pengfei Liu 0003, Ruihao Gong, Dahua Lin, Ningyi Xu
ICML1
2025 Hierachical Balance Packing: Towards Efficient Supervised Fine-tuning for Long-Context LLM
abstract
Training Long-Context Large Language Models (LLMs) is challenging, as hybrid training with long-context and short-context data often leads to workload imbalances. Existing works mainly use data packing to alleviate this issue, but fail to consider imbalanced attention computation and wasted communication overhead. This paper proposes Hierarchical Balance Packing (HBP), which designs a novel batch-construction method and training recipe to address those inefficiencies. In particular, the HBP constructs multi-level data packing groups, each optimized with a distinct packing length. It assigns training samples to their optimal groups and configures each group with the most effective settings, including sequential parallelism degree and gradient checkpointing configuration. To effectively utilize multi-level groups of data, we design a dynamic training pipeline specifically tailored to HBP, including curriculum learning, adaptive sequential parallelism, and stable loss. Our extensive experiments demonstrate that our method significantly reduces training time over multiple datasets and open-source models while maintaining strong performance. For the largest DeepSeek-V2 (236B) MoE model, our method speeds up the training by 2.4$\times$ with competitive performance. Codes will be released at https://github.com/ModelTC/HBP.
Yongqiang Yao, Jingru Tan, Kaihuan Liang, Feizhao Zhang, Yazhe Niu, Ruihao Gong, Dahua Lin, Ningyi Xu
NeurIPS1
2025 Robust long-tailed recognition with distribution-aware adversarial example generation
Bo Li 0126, Yongqiang Yao, Jingru Tan, Dandan Zhu 0001, Ruihao Gong, Ye Luo 0004
Neural Networks2
2024 Towards Frame Rate Agnostic Multi-object Tracking
Lei Bai 0001, Yongqiang Yao, Fengwei Yu, Wanli Ouyang
Int. J. Comput. Vis.3
2024 Rectify representation bias in vision-language models for long-tailed recognition
Bo Li 0126, Yongqiang Yao, Jingru Tan, Ruihao Gong, Ye Luo 0004
Neural Networks2
2024 Similarity- and Quality-Guided Relation Learning for Joint Detection and Tracking
abstract
Joint detection and tracking, which solves two fundamental vision challenges in a unified manner, is a challenging topic in computer vision. In this area, the proper use of spatial-temporal information in videos can help reduce local defects and improve the quality of feature representations. Although modeling low-level (usually pixel-wise) spatial-temporal information has been studied, instance-level spatial-temporal correlations (i.e., relations between semantic regions in which instances have occurred) have not been fully exploited. In comparison, modeling instance-level correlation is a more flexible and reasonable way to enhance feature representations. However, we have found that conventional instance-level relation learning that works for the separate tasks of detection or tracking is not effective in joint tasks in which a variety of scenarios may be presented. To try to resolve this problem, in this study, we effectively exploited instance-level spatial-temporal semantic information for joint detection and tracking via a joint relation learning pipeline with a novel relation learning mechanism called Similarity- and Quality-Guided Attention (SQGA). Specifically, we added task-specific SQGA relation modules before the corresponding task prediction heads to refine the instance feature representation using features of other reference instances in the neighboring frames; these features are aggregated on the basis of relational affinities. In particular, in SQGA, relational affinities were factorized to similarity and quality terms so that fine-grained supervision rules could be applied. Then we added task-specific attention losses for each SQGA relation module, resulting in a better feature aggregation for the corresponding task. Quantitative experiments based on several challenging multi-object tracking benchmarks showed that our approach was more effective than the baselines and provided competitive results compared with recent state-of-the-art methods.
Lei Bai 0001, Yongqiang Yao, Weihao Gan, Wei Wu 0021, Wanli Ouyang
IEEE Trans. Multim.3
2023 Program Translation via Code Distillation
abstract
Software version migration and program translation are an important and costly part of the lifecycle of large codebases.Traditional machine translation relies on parallel corpora for supervised translation, which is not feasible for program translation due to a dearth of aligned data.Recent unsupervised neural machine translation techniques have overcome data limitations by included techniques such as back translation and low level compiler intermediate representations (IR).These methods face significant challenges due to the noise in code snippet alignment and the diversity of IRs respectively.In this paper we propose a novel model called Code Distillation (CoDist) whereby we capture the semantic and structural equivalence of code in a language agnostic intermediate representation.Distilled code serves as a translation pivot for any programming language, leading by construction to parallel corpora which scale to all available source code by simply applying the distillation compiler.We demonstrate that our approach achieves state-of-the-art performance on CodeXGLUE and TransCoder GeeksForGeeks translation benchmarks, with an average absolute increase of 12.7% on the TransCoder GeeksforGeeks translation benchmark compare to TransCoder-ST.
Yufan Huang, Mengnan Qi, Yongqiang Yao, Maoquan Wang, Bin Gu 0001, Colin B. Clement, Neel Sundaresan
EMNLP3
2023 SUT: Active Defects Probing for Transcompiler Models
abstract
Program translation, i.e. transcompilation has been attracting increasing attention from researchers due to its enormous application value.However, we observe that current program translating models still make elementary syntax errors, particularly when the source language uses syntax elements not present in the target language, which is exactly what developers are concerned about while may not be well exposed by frequently used metrics such as BLEU, CodeBLEU and Computation Accuracy.In this paper, we focus on evaluating the model's ability to address these basic syntax errors and developed an novel active defects probing suite, the Syntactic Unit Tests (SUT) and highly interpretable evaluation harness including Syntax Unit Test Accuracy (SUT Acc) metric and Syntax Element Test Score (SETS), to help diagnose and promote progress in this area.Our Syntactic Unit Test fills the gap in the community for a fine-grained evaluation dataset for program translation.Experimental analysis shows that our evaluation harness is more accurate, reliable, and in line with human judgments compared to previous metrics.
Mengnan Qi, Yufan Huang, Maoquan Wang, Yongqiang Yao, Bin Gu 0001, Colin B. Clement, Neel Sundaresan
EMNLP4
2023 The Equalization Losses: Gradient-Driven Training for Long-tailed Object Recognition
abstract
Long-tail distribution is widely spread in real-world applications. Due to the extremely small ratio of instances, tail categories often show inferior accuracy. In this paper, we find such performance bottleneck is mainly caused by the imbalanced gradients, which can be categorized into two parts: (1) positive part, deriving from the samples of the same category, and (2) negative part, contributed by other categories. Based on comprehensive experiments, it is also observed that the gradient ratio of accumulated positives to negatives is a good indicator to measure how balanced a category is trained. Inspired by this, we come up with a gradient-driven training mechanism to tackle the long-tail problem: re-balancing the positive/negative gradients dynamically according to current accumulative gradients, with a unified goal of achieving balance gradient ratios. Taking advantage of the simple and flexible gradient mechanism, we introduce a new family of gradient-driven loss functions, namely equalization losses. We conduct extensive experiments on a wide spectrum of visual tasks, including two-stage/single-stage long-tailed object detection (LVIS), long-tailed image classification (ImageNet-LT, Places-LT, iNaturalist), and long-tailed semantic segmentation (ADE20 K). Our method consistently outperforms the baseline models, demonstrating the effectiveness and generalization ability of the proposed equalization losses.
Jingru Tan, Bo Li 0126, Yongqiang Yao, Fengwei Yu, Tong He 0001, Wanli Ouyang
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 PicassoNet: Searching Adaptive Architecture for Efficient Facial Landmark Localization
abstract
Since recent facial landmark localization methods achieve satisfying accuracy, few of them enable fast inference speed, which, however, is critical in many real-world facial applications. Existing methods typically employ complicated network structure and predict all the key points through uniform computation, which is inefficient since individual facial part might take different computation to obtain the best performance. Taking both accuracy and efficiency into consideration, we propose the PicassoNet, a lightweight cascaded facial landmark detector with adaptive computation for individual facial part. Different from the conventional cascaded methods, PicassoNet integrates refinement submodules into a single network with group convolution, where each convolution group predicts landmarks from an individual facial part. Note that the groups’ structures are flexible in the training process. Then, a novel grouping search algorithm is proposed to optimize the group division. With formulating the optimization as a network architecture search (NAS) problem, the grouping search adaptively allocates computation to each group and obtains an efficient structure. In addition, we propose a boundary-aware loss to optimize along tangent and normal of facial boundaries, instead of optimizing along horizontal and vertical as the conventional loss (L2, SmoothL1, WingLoss, and so on) do. The novel loss improves the joint locations of predicted keypoints. Experiments on three benchmark datasets AFLW, 300W, and WFLW show that the proposed method runs over$6\times $times faster than the state of the arts and meanwhile achieves comparable accuracy.
Tiancheng Wen, Zhonggan Ding, Yongqiang Yao, Yaxiong Wang, Xueming Qian
IEEE Trans. Neural Networks Learn. Syst.3
2022 Equalized Focal Loss for Dense Long-Tailed Object Detection
abstract
Despite the recent success of long-tailed object detection, almost all long-tailed object detectors are developed based on the two-stage paradigm. In practice, one-stage detectors are more prevalent in the industry because they have a simple and fast pipeline that is easy to deploy. However, in the long-tailed scenario, this line of work has not been explored so far. In this paper, we investigate whether one-stage detectors can perform well in this case. We discover the primary obstacle that prevents one-stage detectors from achieving excellent performance is: categories suffer from different degrees of positive-negative imbalance problems under the long-tailed data distribution. The conventional focal loss balances the training process with the same modulating factor for all categories, thus failing to handle the long-tailed problem. To address this issue, we propose the Equalized Focal Loss (EFL) that rebalances the loss contribution of positive and negative samples of different categories independently according to their imbalance degrees. Specifically, EFL adopts a category-relevant modulating factor which can be adjusted dynamically by the training status of different categories. Extensive experiments conducted on the challenging LVIS v1 benchmark demonstrate the effectiveness of our proposed method. With an end-to-end training pipeline, EFL achieves 29.2% in terms of overall AP and obtains significant performance improvements on rare categories, surpassing all existing state-of-the-art methods. The code is available at https: //github.com/ModelTC/EOD.
Bo Li 0126, Yongqiang Yao, Jingru Tan, Fengwei Yu, Ye Luo 0004
CVPR2
2022 Attention Mechanism Based on Improved Spatial-Temporal Convolutional Neural Networks for Traffic Police Gesture Recognition
abstract
Human action recognition has attracted extensive research efforts in recent years, in which traffic police gesture recognition is important for self-driving vehicles. One of the crucial challenges in this task is how to find a representation method based on spatial-temporal features. However, existing methods performed poorly in spatial and temporal information fusion, and how to extract features of traffic police gestures has not been well researched. This paper proposes an attention mechanism based on the improved spatial-temporal convolutional neural network (AMSTCNN) for traffic police gesture recognition. This method focuses on the action part of traffic police and uses the correlation between spatial and temporal features to recognize traffic police gestures, so as to ensure that traffic police gesture information is not lost. Specifically, AMSTCNN integrates spatial and temporal information, uses weight matching to pay more attention to the region where human action occurs, and extracts region proposals of the image. Finally, we use Softmax to classify actions after spatial-temporal feature fusion. AMSTCNN can strongly make use of the spatial-temporal information of videos and select effective features to reduce computation. Experiments on AVA and the Chinese traffic police gesture datasets show that our method is superior to several state-of-the-art methods.
Zhixuan Wu, Nan Ma 0002, Yue Gao 0002, Yongqiang Yao
Int. J. Pattern Recognit. Artif. Intell.6
2020 Bridging the Gap Between Anchor-Based and Anchor-Free Detection via Adaptive Training Sample Selection
abstract
Object detection has been dominated by anchor-based detectors for several years. Recently, anchor-free detectors have become popular due to the proposal of FPN and Focal Loss. In this paper, we first point out that the essential difference between anchor-based and anchor-free detection is actually how to define positive and negative training samples, which leads to the performance gap between them. If they adopt the same definition of positive and negative samples during training, there is no obvious difference in the final performance, no matter regressing from a box or a point. This shows that how to select positive and negative training samples is important for current object detectors. Then, we propose an Adaptive Training Sample Selection (ATSS) to automatically select positive and negative samples according to statistical characteristics of object. It significantly improves the performance of anchor-based and anchor-free detectors and bridges the gap between them. Finally, we discuss the necessity of tiling multiple anchors per location on the image to detect objects. Extensive experiments conducted on MS COCO support our aforementioned analysis and conclusions. With the newly introduced ATSS, we improve state-of-the-art detectors by a large margin to 50.7% AP without introducing any overhead. The code is available at https://github.com/sfzhang15/ATSS.
Cheng Chi 0003, Yongqiang Yao, Zhen Lei 0001, Stan Z. Li
CVPR3
2018 Dense Receptive Field for Object Detection
abstract
Current one-stage single-shot detectors such as DSSD and StairNet based on aggregating context information from multiple scales have shown promising accuracy. However, existing multi-scale context fusion techniques are insufficient for detecting objects of different scales. In this paper, we investigate how to detect different objects with different scales with respect to accuracy-vs-speed trade-off. We propose a novel single-shot based detector, called DRFNet which fuses feature maps with different sizes of the receptive field to boost the detection accuracy. Our final model DRFNet detector unifies comprehensive context information from various receptive fields effectively to enable it to detect objects in different sizes with higher accuracy. Experimental results on PASCAL VOC 2007 benchmark (79.6% mAP, 68 FPS) demonstrate that DRFNet is better than other state-of-the-art one-stage detectors similar to FPN. Code is released at https://github.com/yqyao/DRFNet.
Yongqiang Yao, Zesang Huang, Hongliang Bai
ICPR1
2018 Texture and Geometry Scattering Representation-Based Facial Expression Recognition in 2D+3D Videos
abstract
Facial Expression Recognition (FER) is one of the most important topics in the domain of computer vision and pattern recognition, and it has attracted increasing attention for its scientific challenges and application potentials. In this article, we propose a novel and effective approach to FER using multi-model two-dimensional (2D) and 3D videos, which encodes both static and dynamic clues by scattering convolution network. First, a shape-based detection method is introduced to locate the start and the end of an expression in videos; segment its onset, apex, and offset states; and sample the important frames for emotion analysis. Second, the frames in Apex of 2D videos are represented by scattering, conveying static texture details. Those of 3D videos are processed in a similar way, but to highlight static shape details, several geometric maps in terms of multiple order differential quantities, i.e., Normal Maps and Shape Index Maps, are generated as the input of scattering, instead of original smooth facial surfaces. Third, the average of neighboring samples centred at each key texture frame or shape map in Onset is computed, and the scattering features extracted from all the average samples of 2D and 3D videos are then concatenated to capture dynamic texture and shape cues, respectively. Finally, Multiple Kernel Learning is adopted to combine the features in the 2D and 3D modalities and compute similarities to predict the expression label. Thanks to the scattering descriptor, the proposed approach not only encodes distinct local texture and shape variations of different expressions as by several milestone operators, such as SIFT, HOG, and so on, but also captures subtle information hidden in high frequencies in both channels, which is quite crucial to better distinguish expressions that are easily confused. The validation is conducted on the BU-4DFE and BP-4D databa ses, and the accuracies reached are very competitive, indicating its competency for this issue.
Yongqiang Yao, Di Huang 0001, Yunhong Wang 0001, Liming Chen 0002
ACM Trans. Multim. Comput. Commun. Appl.1