Shijia Huang

dblp:211/5807 · DBLP profile ↗
← Back
16ranked-venue papers
4as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 BiM-Rec: Bi-Stream Mamba for Sequential Recommendation
Ivonne Xu, Yuping Yuan, Xiang Li 0117, Shijia Huang, Ge Zhang 0009
ICIC (4)4
2026 An adaptive rank-based coevolutionary learning particle swarm optimization algorithm for server placement in edge computing
Jian Lü 0002, Shijia Huang, Zhihui He, Feng Wang 0048
Expert Syst. Appl.2
2025 Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding
abstract
The rapid advancement of Multimodal Large Language Models (MLLMs) has significantly impacted various multimodal tasks. However, these models face challenges in tasks that require spatial understanding within 3D environments. Efforts to enhance MLLMs, such as incorporating point cloud features, have been made, yet a considerable gap remains between the models’ learned representations and the inherent complexity of 3D scenes. This discrepancy largely stems from the training of MLLMs on predominantly 2D data, which restricts their effectiveness in comprehending 3D spaces. To address this issue, in this paper, we propose a novel generalist model, i.e., Video-3D LLM, for 3D scene understanding. By treating 3D scenes as dynamic videos and incorporating 3D position encoding into these representations, our Video-3D LLM aligns video representations with real-world spatial contexts more accurately. In addition, we have implemented a maximum coverage sampling technique to optimize the trade-off between computational cost and performance. Extensive experiments demonstrate that our model achieves state-of-the-art performance on several 3D scene understanding benchmarks, including ScanRefer, Multi3DRefer, Scan2Cap, ScanQA, and SQA3D. Our code is available at https://github.com/LaVi-Lab/Video-3D-LLM.
Duo Zheng, Shijia Huang, Liwei Wang 0009
CVPR2
2025 Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors
abstract
Previous research has investigated the application of Multimodal Large Language Models (MLLMs) in understanding 3D scenes by interpreting them as videos. These approaches generally depend on comprehensive 3D data inputs, such as point clouds or reconstructed Bird's-Eye View (BEV) maps. In our research, we advance this field by enhancing the capability of MLLMs to understand and reason in 3D spaces directly from video data, without the need for additional 3D input. We propose a novel and efficient method called the Video-3D Geometry Large Language Model (VG LLM). Our approach utilizes a 3D visual geometry encoder to extract 3D prior information from video sequences. This information is then integrated with visual tokens and input into the MLLM. Extensive experiments have shown that our method has achieved substantial improvements in various tasks related to 3D scene understanding and spatial reasoning, all directly learned from video sources. Impressively, our 4B model, which does not rely on explicit 3D data inputs, achieves competitive results compared to existing state-of-the-art methods, and even surpasses the Gemini-1.5-Pro in the VSI-Bench evaluations.
Duo Zheng, Shijia Huang, Yanyang Li, Liwei Wang 0009
NeurIPS2
2025 Optimal Pricing and Resource Configuration for GPU-accelerated Services in Computing Markets
abstract
GPU-accelerated services accelerate computation-intensive tasks by leveraging the parallel processing power of GPUs but incur high costs. Existing researches lack consideration for GPU’s heterogeneous acceleration efficiency for diverse services. In this work, we study SP’s GPU resource pricing and configuration via GPU virtual computing resource blocks (VCRBs) considering users’ heterogeneous tasks in computing markets. We model the interactions between SP and users as a two-stage Stackelberg game. In Stage I, SP purchases GPU resources and designs the prices and computation capacities of GPU VCRBs to maximize its profit. In Stage II, each user decides whether to adopt a GPU VCRB for task acceleration to minimize its cost. Due to the heterogeneous acceleration efficiency and users’ rational decisions, the uniform GPU VCRB design problem for multiple services is discontinuous and challenging to solve. Hence we first consider the multiple design scheme and derive SP’s optimal GPU VCRB design for each type users, where the price increases and the computation capacity decreases with the acceleration efficiency. Then under the uniform design scheme, we derive the optimal GPU VCRB design for two services and obtain that for multiple services via an efficient two-stage algorithm in quadratic time. Simulation results show that the uniform design scheme lowers the GPU VCRB management complexity and suffers only 12.9% profit loss on average compared to the multiple design scheme.
Shijia Huang
VTC2025-Fall2
2025 A Mutual Supervision Framework for Referring Expression Segmentation and Generation
Shijia Huang, Feng Li 0040, Hao Zhang 0097, Shilong Liu 0004, Lei Zhang 0001, Liwei Wang 0009
Int. J. Comput. Vis.1
2024 Towards Learning a Generalist Model for Embodied Navigation
abstract
Building a generalist agent that can interact with the world is the intriguing target of AI systems, thus spurring the research for embodied navigation, where an agent is required to navigate according to instructions or respond to queries. Despite the major progress attained, previous works primarily focus on task-specific agents and lack gen-eralizability to unseen scenarios. Recently, LLMs have pre-sented remarkable capabilities across various fields, and provided a promising opportunity for embodied navigation. Drawing on this, we propose the first generalist model for embodied navigation, NaviLLM. It adapts LLMs to em-bodied navigation by introducing schema-based instruction. The schema-based instruction flexibly casts various tasks into generation problems, thereby unifying a wide range of tasks. This approach allows us to integrate di-verse data sources from various datasets into the training, equipping NaviLLM with a wide range of capabilities required by embodied navigation. We conduct extensive ex-periments to evaluate the performance and generalizability of our model. The experimental results demonstrate that our unified model achieves state-of-the-art performance on CVDN, SOON, and ScanQA. Specifically, it surpasses the previous stats-of-the-art method by a significant margin of 29% in goal progress on CVDN. Moreover, our model also demonstrates strong generalizability and presents im-pressive results on unseen tasks, e.g. embodied question answering and 3D captioning. Our code is available at https://github.com/LaVi-Lab/NaviLLM.
Duo Zheng, Shijia Huang, Lin Zhao 0016, Yiwu Zhong, Liwei Wang 0009
CVPR2
2024 LLaVA-Grounding: Grounded Visual Chat with Large Multimodal Models
Hao Zhang 0097, Hongyang Li 0003, Feng Li 0040, Tianhe Ren, Xueyan Zou, Shilong Liu 0004, Shijia Huang, Jianfeng Gao 0001, Leizhang, Chunyuan Li, Jainwei Yang
ECCV (43)7
2024 A coevolutionary estimation of distribution algorithm based on dynamic differential grouping for mixed-variable optimization problems
Shijia Huang, Feng Wang 0048
Expert Syst. Appl.1
2023 DQ-DETR: Dual Query Detection Transformer for Phrase Extraction and Grounding
abstract
In this paper, we study the problem of visual grounding by considering both phrase extraction and grounding (PEG). In contrast to the previous phrase-known-at-test setting, PEG requires a model to extract phrases from text and locate objects from image simultaneously, which is a more practical setting in real applications. As phrase extraction can be regarded as a 1D text segmentation problem, we formulate PEG as a dual detection problem and propose a novel DQ-DETR model, which introduces dual queries to probe different features from image and text for object prediction and phrase mask prediction. Each pair of dual queries are designed to have shared positional parts but different content parts. Such a design effectively alleviates the difficulty of modality alignment between image and text (in contrast to a single query design) and empowers Transformer decoder to leverage phrase mask-guided attention to improve the performance. To evaluate the performance of PEG, we also propose a new metric CMAP (cross-modal average precision), analogous to the AP metric in object detection. The new metric overcomes the ambiguity of Recall@1 in many-box-to-one-phrase cases in phrase grounding. As a result, our PEG pre-trained DQ-DETR establishes new state-of-the-art results on all visual grounding benchmarks with a ResNet-101 backbone. For example, it achieves 91.04% and 83.51% in terms of recall rate on RefCOCO testA and testB with a ResNet-101 backbone.
Shilong Liu 0004, Shijia Huang, Feng Li 0040, Hao Zhang 0097, Yaoyuan Liang, Hang Su 0006, Jun Zhu 0001, Lei Zhang 0001
AAAI2
2023 MP-Former: Mask-Piloted Transformer for Image Segmentation
abstract
We present a mask-piloted Transformer which improves masked-attention in Mask2Former for image segmentation. The improvement is based on our observation that Mask2Former suffers from inconsistent mask predictions between consecutive decoder layers, which leads to inconsistent optimization goals and low utilization of decoder queries. To address this problem, we propose a mask-piloted training approach, which additionally feeds noised ground-truth masks in masked-attention and trains the model to reconstruct the original ones. Compared with the predicted masks used in mask-attention, the ground-truth masks serve as a pilot and effectively alleviate the negative impact of inaccurate mask predictions in Mask2Former. Based on this technique, our MP-Former achieves a remarkable performance improvement on all three image segmentation tasks (instance, panoptic, and semantic), yielding +2.3AP and +1.6mIoU on the Cityscapes instance and semantic segmentation tasks with a ResNet-50 backbone. Our method also significantly speeds up the training, outperforming Mask2Former with half of the number of training epochs on ADE20K with both a ResNet-50 and a Swin-L backbones. Moreover, our method only introduces little computation during training and no extra computation during inference. Our code will be released at https://github.com/IDEA-Research/MP-Former.
Hao Zhang 0097, Feng Li 0040, Huaizhe Xu, Shijia Huang, Shilong Liu 0004, Lionel M. Ni, Lei Zhang 0001
CVPR4
2023 Learning Preference Model for LLMs via Automatic Preference Data Generation
abstract
Despite the advanced capacities of the stateof-the-art large language models (LLMs), they suffer from issues of hallucination, stereotype, etc. Preference models play an important role in LLM alignment, yet training preference models predominantly rely on human-annotated data.This reliance limits their versatility and scalability.In this paper, we propose learning the preference model for LLMs via automatic preference data generation (AutoPM).Our approach involves both In-Breadth Data Generation, which elicits pairwise preference data from LLMs following the helpful-honestharmless (HHH) criteria, and In-Depth Data Generation, which enriches the dataset with responses spanning a wide quality range.With HHH-guided preference data, our approach simultaneously enables the LLMs to learn human preferences and align with human values.Quantitative assessments on five benchmark datasets demonstrate the reliability and potential of AutoPM, pointing out a more general and scalable way to improve LLM performance.
Shijia Huang, Jianqiao Zhao, Yanyang Li, Liwei Wang 0009
EMNLP1
2023 DSGN++: Exploiting Visual-Spatial Relation for Stereo-Based 3D Detectors
abstract
Camera-based 3D object detectors are welcome due to their wider deployment and lower price than LiDAR sensors. We first revisit the prior stereo detector DSGN for its stereo volume construction ways for representing both 3D geometry and semantics. We polish the stereo modeling and propose the advanced version, DSGN++, aiming to enhance effective information flow throughout the 2D-to-3D pipeline in three main aspects. First, to effectively lift the 2D information to stereo volume, we propose depth-wise plane sweeping (DPS) that allows denser connections and extracts depth-guided features. Second, for grasping differently spaced features, we present a novel stereo volume - Dual-view Stereo Volume (DSV) that integrates front-view and top-view features and reconstructs sub-voxel depth in the camera frustum. Third, as the foreground region becomes less dominant in 3D space, we propose a multi-modal data editing strategy - Stereo-LiDAR Copy-Paste, which ensures cross-modal alignment and improves data efficiency. Without bells and whistles, extensive experiments in various modality setups on the popular KITTI benchmark show that our method consistently outperforms other camera-based 3D detectors for all categories. Code is available at https://github.com/chenyilun95/DSGN2.
Shijia Huang, Shu Liu 0005, Bei Yu 0001, Jiaya Jia
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Multi-View Transformer for 3D Visual Grounding
abstract
The 3D visual grounding task aims to ground a natural language description to the targeted object in a 3D scene, which is usually represented in 3D point clouds. Previous works studied visual grounding under specific views. The vision-language correspondence learned by this way can easily fail once the view changes. In this paper, we propose a Multi-View Transformer (MVT) for 3D visual grounding. We project the 3D scene to a multi-view space, in which the position information of the 3D scene under different views are modeled simultaneously and aggregated together. The multi-view space enables the network to learn a more robust multi-modal representation for 3D visual grounding and eliminates the dependence on specific views. Extensive experiments show that our approach significantly outperforms all state-of-the-art methods. Specifically, on Nr3D and Sr3D datasets, our method outperforms the best competitor by 11.2% and 7.1% and even surpasses recent work with extra 2D assistance by 5.9% and 6.6%. Our code is available at https://github.com/sega-hsj/MVT-3DVG.
Shijia Huang, Jiaya Jia, Liwei Wang 0009
CVPR1
2018 Surrogate-Assisted Multi-Tasking Memetic Algorithm
abstract
This paper proposes a surrogate-assisted multitasking memetic algorithm (SaM-MA) for multi-tasking optimization. In the proposed SaM-MA, the population is divided into multiple sub-populations, with each sub-population focusing on solving one task. Each sub-population is evolved by three components. The first is the global search component which used differential evolutionary algorithm to search for the global optimal solution for the corresponding task. The second component is a surrogate model with Gaussian process, which is used to predict the best solution, so as to reduce the number of fitness evaluations and to improve the search efficiency. The third component is the local search component which utilizes the CMA-ES to locally exploiting the neighboring regions of promising solutions. In addition, the crossover operators in the global search component are extended so as to facilitate knowledge transfer between sub-populations, The proposed SaM-MA is tested on nine benchmark multi-tasking optimization problems in the CEC2017 competition. The experiment results have demonstrated the efficacy of the proposed SaM-MA in terms of solution accuracy and search efficiency.
Dingnan Liu, Shijia Huang, Jinghui Zhong
CEC2
2017 Pedestrian Detection via Structure-Sensitive Deep Representation Learning
Deliang Huang, Shijia Huang, Hefeng Wu
ICIG (1)2