Yuqi Huo

dblp:219/6931 · DBLP profile ↗
← Back
20ranked-venue papers
5as first author
13since 2021 · last 2026
0009-0009-1202-4791ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 5 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 AITQE: An Adaptive Image-Text Quality Enhancer for Scalable MLLM Pretraining
abstract
Multimodal large language models (MLLMs) have made significant strides by integrating visual and textual modalities. A critical factor in training MLLMs is the quality of image-text pairs within multimodal pretraining datasets. However, in the process of high-quality data curation, filter-based paradigms often discard a substantial portion of high-quality images due to inadequate semantic alignment between images and texts, leading to inefficiency in data utilization and scalability. In this paper, we propose the Adaptive Image-Text Quality Enhancer (AITQE), a model that dynamically assesses and enhances the quality of image-text pairs. AITQE employs a text rewriting mechanism for low-quality pairs and incorporates a negative sample learning strategy to improve evaluative capabilities by integrating deliberately generated low-quality samples during training. Unlike prior approaches that significantly alter text distributions, our method minimally adjusts text to preserve data volume while enhancing quality. Experimental results demonstrate that AITQE surpasses existing methods on various benchmarks, effectively leveraging raw data and scaling with increasing data volumes. Codes and model are available at https://github.com/hanhuang22/AITQE.
Yuqi Huo, Zijia Zhao, Haoyu Lu, Bingning Wang, Qiang Liu 0006, Weipeng Chen, Liang Wang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 Efficient Motion-Aware Video MLLM
abstract
Most current video MLLMs rely on uniform frame sampling and image-level encoders, resulting in inefficient data processing and limited motion awareness. To address these challenges, we introduce EMA, an Efficient Motion-Aware video MLLM that utilizes compressed video structures as inputs. We propose a motion-aware GOP (Group of Pictures) encoder that fuses spatial and motion information within a GOP unit in the compressed video stream, generating compact, informative visual tokens. By integrating fewer but denser RGB frames with more but sparser motion vectors in this native slow-fast input architecture, our approach reduces redundancy and enhances motion representation. Additionally, we introduce MotionBench, a benchmark for evaluating motion understanding across four motion types: linear, curved, rotational, and contact-based. Experimental results show that EMA achieves state-of-the-art performance on both MotionBench and popular video question answering benchmarks, while reducing inference costs. Moreover, EMA demonstrates strong scalability, as evidenced by its competitive performance on long video understanding benchmarks.
Zijia Zhao, Yuqi Huo, Tongtian Yue, Longteng Guo, Haoyu Lu, Bingning Wang, Weipeng Chen, Jing Liu 0001
CVPR2
2025 Exploring the Design Space of Visual Context Representation in Video MLLMs
abstract
Video Multimodal Large Language Models~(MLLMs) have shown remarkable capability of understanding the video semantics on various downstream tasks. Despite the advancements, there is still a lack of systematic research on visual context representation, which refers to the scheme to select frames from a video and further select the tokens from a frame. In this paper, we explore the design space for visual context representation, and aim to improve the performance of video MLLMs by finding more effective representation schemes. Firstly, we formulate the task of visual context representation as a constrained optimization problem, and model the language modeling loss as a function of the number of frames and the number of embeddings (or tokens) per frame, given the maximum visual context window size. Then, we explore the scaling effects in frame selection and token selection respectively, and fit the corresponding function curve by conducting extensive empirical experiments. We examine the effectiveness of typical selection strategies and present empirical findings to determine the two factors. Furthermore, we study the joint effect of frame selection and token selection, and derive the optimal formula for determining the two factors. We demonstrate that the derived optimal settings show alignment with the best-performed results of empirical experiments. The data and code are available at: https://github.com/RUCAIBox/Opt-Visor.
Yifan Du 0002, Yuqi Huo, Kun Zhou 0002, Zijia Zhao, Haoyu Lu, Wayne Xin Zhao, Bingning Wang, Weipeng Chen, Ji-Rong Wen
ICLR2
2025 Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs
abstract
Video understanding is a crucial next step for multimodal large language models (MLLMs). Various benchmarks are introduced for better evaluating the MLLMs. Nevertheless, current video benchmarks are still inefficient for evaluating video models during iterative development due to the high cost of constructing datasets and the difficulty in isolating specific skills. In this paper, we propose VideoNIAH (Video Needle in A Haystack), a benchmark construction framework through synthetic video generation. VideoNIAH decouples video content from their query-responses by inserting unrelated visual 'needles' into original videos. The framework automates the generation of query-response pairs using predefined rules, minimizing manual labor. The queries focus on specific aspects of video understanding, enabling more skill-specific evaluations. The separation between video content and the queries also allow for increased video variety and evaluations across different lengths. Utilizing VideoNIAH, we compile a video benchmark, VNBench, which includes tasks such as retrieval, ordering, and counting to evaluate three key aspects of video understanding: temporal perception, chronological ordering, and spatio-temporal coherence. We conduct a comprehensive evaluation of both proprietary and open-source models, uncovering significant differences in their video understanding capabilities across various tasks. Additionally, we perform an in-depth analysis of the test results and model configurations. Based on these findings, we provide some advice for improving video MLLM training, offering valuable insights to guide future research and model development.
Zijia Zhao, Haoyu Lu, Yuqi Huo, Yifan Du 0002, Tongtian Yue, Longteng Guo, Bingning Wang, Weipeng Chen, Jing Liu 0001
ICLR3
2025 NGAP Feature Fusion Hybrid Network Attack Detection for 5G Edge Security
abstract
The fifth-generation (5G) mobile network is a critical infrastructure for cellular communication, requiring the confidentiality, integrity and availability of services. However, the inherent vulnerabilities of Radio Access Network (RAN) allow the attacker to exploit vulnerabilities in 3GPP specifications or implementation to compromise user privacy and disrupt services. Existing defense methods are limited by the reliance on manual analysis and rule-based detection, which fails to detect novel and evolving threats. We propose NGAPAD, the first system designed to automatically monitor and analyze 5G edge attack based on Next Generation Application Protocol (NGAP). NGAPAD provides a feasible solution to overcome challenges of threat pattern universality, protocol specificity and data efficiency in 5G edge security. We design a new NGAP telemetry format and a dual-branch hybrid network to achieve precise and efficient attack detection. We constructed a high-quality dataset and evaluated it experimentally on 5G simulation network. NGAPAD achieved the optimal performance metrics by sequence length tuning, achieving 99.31% F1 Score with the length of 12. The system successfully detected 18 out of 22 known edge attacks and achieved 98.3% Accuracy against unknown attacks generated by fuzzing of NGAP protocol.
Shaocong Feng, Baojiang Cui, Shengjia Chang, Yuqi Huo
IEEE Internet Things J.5
2024 UniAdapter: Unified Parameter-Efficient Transfer Learning for Cross-modal Modeling
abstract
Large-scale vision-language pre-trained models have shown promising transferability to various downstream tasks. As the size of these foundation models and the number of downstream tasks grow, the standard full fine-tuning paradigm becomes unsustainable due to heavy computational and storage costs. This paper proposes UniAdapter, which unifies unimodal and multimodal adapters for parameter-efficient cross-modal adaptation on pre-trained vision-language models. Specifically, adapters are distributed to different modalities and their interactions, with the total number of tunable parameters reduced by partial weight sharing. The unified and knowledge-sharing design enables powerful cross-modal representations that can benefit various downstream tasks, requiring only 1.0%-2.0% tunable parameters of the pre-trained model. Extensive experiments on 7 cross-modal downstream benchmarks (including video-text retrieval, image-text retrieval, VideoQA, VQA and Caption) show that in most cases, UniAdapter not only outperforms the state-of-the-arts, but even beats the full fine-tuning strategy. Particularly, on the MSRVTT retrieval task, UniAdapter achieves 49.7% recall@1 with 2.2% model parameters, outperforming the latest competitors by 2.0%. The code and models are available at https://github.com/RERV/UniAdapter.
Haoyu Lu, Yuqi Huo, Guoxing Yang, Zhiwu Lu 0001, Masayoshi Tomizuka, Mingyu Ding
ICLR2
2024 VDT: General-purpose Video Diffusion Transformers via Mask Modeling
abstract
This work introduces Video Diffusion Transformer (VDT), which pioneers the use of transformers in diffusion-based video generation. It features transformer blocks with modularized temporal and spatial attention modules to leverage the rich spatial-temporal representation inherited in transformers. Additionally, we propose a unified spatial-temporal mask modeling mechanism, seamlessly integrated with the model, to cater to diverse video generation scenarios. VDT offers several appealing benefits. (1) It excels at capturing temporal dependencies to produce temporally consistent video frames and even simulate the physics and dynamics of 3D objects over time. (2) It facilitates flexible conditioning information, e.g., simple concatenation in the token space, effectively unifying different token lengths and modalities. (3) Pairing with our proposed spatial-temporal mask modeling mechanism, it becomes a general-purpose video diffuser for harnessing a range of tasks, including unconditional generation, video prediction, interpolation, animation, and completion, etc. Extensive experiments on these tasks spanning various scenarios, including autonomous driving, natural weather, human action, and physics-based simulation, demonstrate the effectiveness of VDT. Moreover, we provide a comprehensive study on the capabilities of VDT in capturing accurate temporal dependencies, handling conditioning information, and the spatial-temporal mask modeling mechanism. Additionally, we present comprehensive studies on how VDT handles conditioning information with the mask modeling mechanism, which we believe will benefit future research and advance the field. Codes and models are available at the https://VDT-2023.github.io.
Haoyu Lu, Guoxing Yang, Nanyi Fei, Yuqi Huo, Zhiwu Lu 0001, Ping Luo 0002, Mingyu Ding
ICLR4
2022 COTS: Collaborative Two-Stream Vision-Language Pre-Training Model for Cross-Modal Retrieval
abstract
Large-scale single-stream pre-training has shown dramatic performance in image-text retrieval. Regrettably, it faces low inference efficiency due to heavy attention layers. Recently, two-stream methods like CLIP and ALIGN with high inference efficiency have also shown promising performance, however, they only consider instance-level alignment between the two streams (thus there is still room for improvement). To overcome these limitations, we propose a novel COllaborative Two-Stream vision-language pretraining model termed COTS for image-text retrieval by enhancing cross-modal interaction. In addition to instance-level alignment via momentum contrastive learning, we leverage two extra levels of cross-modal interactions in our COTS: (1) Token-level interaction - a masked vision-language modeling (MVLM) learning objective is devised without using a cross-stream network module, where variational autoencoder is imposed on the visual encoder to generate visual tokens for each image. (2) Task-level interaction - a KL-alignment learning objective is devised between text-to-image and image-to-text retrieval tasks, where the probability distribution per task is computed with the negative queues in momentum contrastive learning. Under a fair comparison setting, our COTS achieves the highest performance among all two-stream methods and comparable performance (but with 10,800× faster in inference) w.r.t. the latest single-stream methods. Importantly, our COTS is also applicable to text-to-video retrieval, yielding new state-of-the-art on the widely-used MSR-VTT dataset.
Haoyu Lu, Nanyi Fei, Yuqi Huo, Yizhao Gao 0004, Zhiwu Lu 0001, Ji-Rong Wen
CVPR3
2022 Learning Versatile Neural Architectures by Propagating Network Codes
Mingyu Ding, Yuqi Huo, Haoyu Lu, Zhe Wang 0006, Zhiwu Lu 0001, Jingdong Wang 0001, Ping Luo 0002
ICLR2
2022 LGDN: Language-Guided Denoising Network for Video-Language Modeling
abstract
Video-language modeling has attracted much attention with the rapid growth of web videos. Most existing methods assume that the video frames and text description are semantically correlated, and focus on video-language modeling at video level. However, this hypothesis often fails for two reasons: (1) With the rich semantics of video contents, it is difficult to cover all frames with a single video-level description; (2) A raw video typically has noisy/meaningless information (e.g., scenery shot, transition or teaser). Although a number of recent works deploy attention mechanism to alleviate this problem, the irrelevant/noisy information still makes it very difficult to address. To overcome such challenge, we thus propose an efficient and effective model, termed Language-Guided Denoising Network (LGDN), for video-language modeling. Different from most existing methods that utilize all extracted video frames, LGDN dynamically filters out the misaligned or redundant frames under the language supervision and obtains only 2--4 salient frames per video for cross-modal token-level alignment. Extensive experiments on five public datasets show that our LGDN outperforms the state-of-the-arts by large margins. We also provide detailed ablation study to reveal the critical importance of solving the noise issue, in hope of inspiring future video-language work.
Haoyu Lu, Mingyu Ding, Nanyi Fei, Yuqi Huo, Zhiwu Lu 0001
NeurIPS4
2021 Complex Action Segmentation in Compressed Videos
abstract
Complex action segmentation aims to detect what actions and when they happen in fine-grained level from long videos. Despite the fact that videos are often stored in a compressed format (e.g., MPEG-4), most existing approaches are proposed to directly model raw RGB videos: when only compressed videos are accessible, they have to first decode these videos, which is very time-consuming. In this paper, by explicitly leveraging the ‘compressed’ characteristic of compressed videos, we are the first to address the challenging task of complex action segmentation in compressed videos. To extract meaningful representations for complex action segmentation, we introduce the GOP-Level Compressed features (Golec), which can be obtained directly from compressed videos without video decompression. Importantly, by taking GOPs as the atomic units of actions, our Golec representation is intrinsically suitable for fine-grained action segmentation. Moreover, to remedy the coarser motion vectors (compared with optical flows which are computed from raw frames) used in our Golec representation for capturing the temporal context, we propose a new Bi-path knowledge distillation strategy. Extensive experiments show the effectiveness of our Golec representation and the Bi-path strategy. Importantly, our proposed model for complex action detection not only runs 5.2 times faster but also achieves significantly better results than the state-of-the-art alternatives using raw videos.
Hongfeng Han, Guoxing Yang, Yuqi Huo, Zhiwu Lu 0001, Ji-Rong Wen
ICME3
2021 Self-Supervised Video Representation Learning with Constrained Spatiotemporal Jigsaw
abstract
This paper proposes a novel pretext task for self-supervised video representation learning by exploiting spatiotemporal continuity in videos. It is motivated by the fact that videos are spatiotemporal by nature and a representation learned by detecting spatiotemporal continuity/discontinuity is thus beneficial for downstream video content analysis tasks. A natural choice of such a pretext task is to construct spatiotemporal (3D) jigsaw puzzles and learn to solve them. However, as we demonstrate in the experiments, this task turns out to be intractable. We thus propose Constrained Spatiotemporal Jigsaw (CSJ) whereby the 3D jigsaws are formed in a constrained manner to ensure that large continuous spatiotemporal cuboids exist. This provides sufficient cues for the model to reason about the continuity. Instead of solving them directly, which could still be extremely hard, we carefully design four surrogate tasks that are more solvable. The four tasks aim to learn representations sensitive to spatiotemporal continuity at both the local and global levels. Extensive experiments show that our CSJ achieves state-of-the-art on various benchmarks.
Yuqi Huo, Mingyu Ding, Haoyu Lu, Mingqian Tang, Zhiwu Lu 0001, Tao Xiang 0002
IJCAI1
2021 Compressed Video Contrastive Learning
abstract
This work concerns self-supervised video representation learning (SSVRL), one topic that has received much attention recently. Since videos are storage-intensive and contain a rich source of visual content, models designed for SSVRL are expected to be storage- and computation-efficient, as well as effective. However, most existing methods only focus on one of the two objectives, failing to consider both at the same time. In this work, for the first time, the seemingly contradictory goals are simultaneously achieved by exploiting compressed videos and capturing mutual information between two input streams. Specifically, a novel Motion Vector based Cross Guidance Contrastive learning approach (MVCGC) is proposed. For storage and computation efficiency, we choose to directly decode RGB frames and motion vectors (that resemble low-resolution optical flows) from compressed videos on-the-fly. To enhance the representation ability of the motion vectors, hence the effectiveness of our method, we design a cross guidance contrastive learning algorithm based on multi-instance InfoNCE loss, where motion vectors can take supervision signals from RGB frames and vice versa. Comprehensive experiments on two downstream tasks show that our MVCGC yields new state-of-the-art while being significantly more efficient than its competitors.
Yuqi Huo, Mingyu Ding, Haoyu Lu, Nanyi Fei, Zhiwu Lu 0001, Ji-Rong Wen, Ping Luo 0002
NeurIPS1
2020 Learning Depth-Guided Convolutions for Monocular 3D Object Detection
abstract
3D object detection from a single image without LiDAR is a challenging task due to the lack of accurate depth information. Conventional 2D convolutions are unsuitable for this task because they fail to capture local object and its scale information, which are vital for 3D object detection. To better represent 3D structure, prior arts typically transform depth maps estimated from 2D images into a pseudo-LiDAR representation, and then apply existing 3D point-cloud based object detectors. However, their results depend heavily on the accuracy of the estimated depth maps, resulting in suboptimal performance. In this work, instead of using pseudo-LiDAR representation, we improve the fundamental 2D fully convolutions by proposing a new local convolutional network (LCN), termed Depth-guided Dynamic-Depthwise-Dilated LCN (D4LCN), where the filters and their receptive fields can be automatically learned from image-based depth maps, making different pixels of different images have different filters. D4LCN overcomes the limitation of conventional 2D convolutions and narrows the gap between image representation and 3D point cloud representation. Extensive experiments show that D4LCN outperforms existing works by large margins. For example, the relative improvement of D4LCN against the state-of-the-art on KITTI is 9.1\% in the moderate setting. D4LCN ranks 1st on KITTI monocular 3D object detection benchmark at the time of submission (car, December 2019). The code is available at https://github.com/dingmyu/D4LCN
Mingyu Ding, Yuqi Huo, Hongwei Yi, Zhe Wang 0006, Jianping Shi, Zhiwu Lu 0001, Ping Luo 0002
CVPR2
2019 Zero-Shot Learning with Few Seen Class Samples
abstract
Zero-shot learning (ZSL) is originally designed to address the small sample size problem often encountered in computer vision by recognizing unseen object classes without any training samples. Existing ZSL models (particularly deep ones) often assume that hundreds of labelled samples are collected from each seen class. In real-world applications, this assumption tends to become invalid. Therefore, a new ZSL setting is concerned in this paper: each seen class only has few labelled samples, while each unseen class still has no samples. This is more challenging yet more useful/practical than the conventional ZSL setting. To overcome the extreme label scarcity, we choose to obtain more training samples from image search engine for data augmentation: the name of each seen class is used as the query of Google, and the top returned images can be viewed as the noisy labelled samples for this seen class. With the augmented but noisy labelled training data, a novel inductive ZSL model is proposed by formulating label noise reduction (LNR) and semantic projection learning (SPL) within a unified framework: (1) LNR aims to refine the noisy labelled samples for projection learning; (2) SPL aims to learn the projection function with the refined training data. Extensive experiments show that our ZSL model outperforms the state-of-the-art alternatives.
Yuqi Huo, Jiechao Guan, Manli Zhang, Ji-Rong Wen, Zhiwu Lu 0001
ICME1
2019 Coarse-to-Fine Grained Classification
abstract
Fine-grained image classification and retrieval become topical in both computer vision and information retrieval. In real-life scenarios, fine-grained tasks tend to appear along with coarse-grained tasks when the observed object is coming closer. However, in previous works, the combination of fine-grained and coarse-grained tasks was often ignored. In this paper, we define a new problem called coarse-to-fine grained classification (C2FGC) which aims to recognize the classes of objects in multiple resolutions (from low to high). To solve this problem, we propose a novel Multi-linear Pooling with Hierarchy (MLPH) model. Specifically, we first design a multi-linear pooling module to include both trilinear and bilinear pooling, and then formulate the coarse-grained and fine-grained tasks within a unified framework. Experiments on two benchmark datasets show that our model achieves state-of-the-art results.
Yuqi Huo, Yulei Niu, Zhiwu Lu 0001, Ji-Rong Wen
SIGIR1
2018 DeepInsight: Multi-Task Multi-Scale Deep Learning for Mental Disorder Diagnosis
Mingyu Ding, Yuqi Huo, Zhiwu Lu 0001
BMVC2
2018 Zero-Shot Learning with Superclasses
Yuqi Huo, Mingyu Ding, Ji-Rong Wen, Zhiwu Lu 0001
ICONIP (3)1
2018 InsightGAN: Semi-Supervised Feature Learning with Generative Adversarial Network for Drug Abuse Detection
Guangzhen Liu, Mingyu Ding, Yuqi Huo, Zhiwu Lu 0001
ICONIP (3)5
2018 Not all bug reopens are negative: A case study on eclipse bug reports
Qing Mi, Jacky W. Keung, Yuqi Huo, Solomon Mensah
Inf. Softw. Technol.3