Jialian Wu

dblp:272/1065 · DBLP profile ↗
← Back
16ranked-venue papers
8as first author
13since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 7 first-author · 8 since 2021Artificial intelligence and machine learning · 9 · 4 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Self-Taught Agentic Long Context Understanding
abstract
Yufan Zhuang, Xiaodong Yu, Jialian Wu, Ximeng Sun, Ze Wang, Jiang Liu, Yusheng Su, Jingbo Shang, Zicheng Liu, Emad Barsoum. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yufan Zhuang, Jialian Wu, Ximeng Sun, Ze Wang 0008, Jiang Liu 0014, Yusheng Su, Jingbo Shang, Zicheng Liu 0001, Emad Barsoum
ACL (1)3
2025 TTT-Bench: A Benchmark for Evaluating Reasoning Ability with Simple and Novel Tic-Tac-Toe-style Games
abstract
Large reasoning models (LRMs) have demonstrated impressive reasoning capabilities across a broad range of tasks including Olympiadlevel mathematical problems, indicating evidence of their complex reasoning abilities.While many reasoning benchmarks focus on the STEM domain, the ability of LRMs to reason correctly in broader task domains remains underexplored.In this work, we introduce TTT-Bench, a new benchmark that is designed to evaluate basic strategic, spatial, and logical reasoning abilities in LRMs through a suite of four two-player Tic-Tac-Toe-style games that humans can effortlessly solve from a young age.We propose a simple yet scalable programmatic approach for generating verifiable two-player game problems for TTT-Bench.Although these games are trivial for humans, they require reasoning about the intentions of the opponent, as well as the game board's spatial configurations, to ensure a win.We evaluate a diverse set of state-of-the-art LRMs, and discover that the models that excel at hard math problems frequently fail at these simple reasoning games.Further testing reveals that our evaluated reasoning models score on average ↓ 41% & ↓ 5% lower on TTT-Bench compared to MATH 500 & AIME 2024 respectively, with larger models achieving higher performance using shorter reasoning traces, where most of the models struggle on long-term strategic reasoning situations on simple and new TTT-Bench tasks.
Prakamya Mishra, Jiang Liu 0014, Jialian Wu, Zicheng Liu 0001, Emad Barsoum
EMNLP3
2025 Unleashing Hour-Scale Video Training for Long Video-Language Understanding
abstract
Recent long-form video-language understanding benchmarks have driven progress in video large multimodal models (Video-LMMs). However, the scarcity of well-annotated long videos has left the training of hour-long Video-LMMs underexplored. To close this gap, we present VideoMarathon, a large-scale hour-long video instruction-following dataset. This dataset includes around 9,700 hours of long videos sourced from diverse domains, ranging from 3 to 60 minutes per video. Specifically, it contains 3.3M high-quality QA pairs, spanning six fundamental topics: temporality, spatiality, object, action, scene, and event. Compared to existing video instruction datasets, VideoMarathon significantly extends training video durations up to 1 hour, and supports 22 diverse tasks requiring both short- and long-term video comprehension. Building on VideoMarathon, we propose Hour-LLaVA, a powerful and efficient Video-LMM for hour-scale video-language modeling. It enables hour-long video training and inference at 1-FPS sampling by leveraging a memory augmentation module, which adaptively integrates question-relevant and spatiotemporally informative semantics from the cached full video context. In our experiments, Hour-LLaVA achieves the best performance on multiple representative long video-language benchmarks, demonstrating the high quality of the VideoMarathon dataset and the superiority of the Hour-LLaVA model.
Jialian Wu, Ximeng Sun, Ze Wang 0008, Jiang Liu 0014, Yusheng Su, Hao Chen 0102, Jiebo Luo 0001, Zicheng Liu 0001, Emad Barsoum
NeurIPS2
2024 GRiT: A Generative Region-to-Text Transformer for Object Understanding
Jialian Wu, Zhengyuan Yang, Zhe Gan, Zicheng Liu 0001, Junsong Yuan 0001
ECCV (80)1
2024 Improved Group Sparse Modal Decomposition Methods With Applications to Fault Diagnosis of Rotating Machinery
abstract
Group-sparse mode decomposition (GSMD) is an efficient signal decomposition algorithm for separating harmonic signals, but fails to split modes for periodic pulse signals. In order to address the limitations of the GSMD algorithm in dealing with periodic pulse signals, this study proposes two improved GSMD methods, namely adaptive adjusted bandwidth sparse mode decomposition (AABSMD) and adaptive Gaussian window sparse mode decomposition (AGWSMD). The AABSMD method utilizes an iterative least-squares curve fitting approach to plot the energy spectrum and adjusts the filter using a –3 dB bandwidth, which avoids unreasonable bandwidth estimation. The AGWSMD method employs a segmented quantile regression method to fit the energy spectrum and utilizes a Gaussian window as a filter to extract the signal, which selects energy entropy as the parameter to optimize the Gaussian function. Both methods overcome the defects of splitting periodic pulse signals in GSMD and provide more accurate determination of the number of modes, resulting in a well-performed decomposition. In addition, the AGWSMD method exhibits outstanding reconstructive performance and high efficiency. Numerical and experimental results indicate the effectiveness and superiority of the proposed methods, which can be successfully applied to fault diagnosis of rotating machineries.
Yueyang Li 0001, Jialian Wu, Dong Zhao 0004
IEEE Trans. Ind. Informatics2
2022 Efficient Video Instance Segmentation via Tracklet Query and Proposal
abstract
Video Instance Segmentation (VIS) aims to simultaneously classify, segment, and track multiple object instances in videos. Recent clip-level VIS takes a short video clip as input each time showing stronger performance than frame-level VIS (tracking-by-segmentation), as more temporal context from multiple frames is utilized. Yet, most clip-level methods are neither end-to-end learnable nor real-time. These limitations are addressed by the recent VIS transformer (VisTR) [25] which performs VIS end-to-end within a clip. However, VisTR suffers from long training time due to its frame-wise dense attention. In addition, VisTR is not fully end-to-end learnable in multiple video clips as it requires a hand-crafted data association to link instance tracklets between successive clips. This paper proposes EfficientVIS, a fully end-to-end framework with efficient training and inference. At the core are tracklet query and tracklet proposal that associate and segment regions-of-interest (RoIs) across space and time by an iterative query-video interaction. We further propose a correspondence learning that makes tracklets linking between clips end-to-end learnable. Compared to VisTR, EfficientVIS requires$15\times$fewer training epochs while achieving state-of-the-art accuracy on the YouTube-VIS benchmark. Meanwhile, our method enables whole video instance segmentation in a single end-to-end pass without data association at all.
Jialian Wu, Sudhir Yarram, Hui Liang 0003, Junsong Yuan 0001, Jayan Eledath, Gérard G. Medioni
CVPR1
2022 Deformable VisTR: Spatio Temporal Deformable Attention for Video Instance Segmentation
abstract
Video instance segmentation (VIS) task requires classifying, segmenting, and tracking object instances over all frames in a video clip. Recently, VisTR [1] has been proposed as end-to-end transformer-based VIS framework, while demonstrating state-of-the-art performance. However, VisTR is slow to converge during training, requiring around 1000 GPU hours due to the high computational cost of its transformer attention module. To improve the training efficiency, we propose Deformable VisTR, leveraging spatio-temporal deformable attention module that only attends to a small fixed set of key spatio-temporal sampling points around a reference point. This enables Deformable VisTR to achieve linear computation in the size of spatio-temporal feature maps. Moreover, it can achieve on par performance as the original VisTR with 10× less GPU training hours. We validate the effectiveness of our method on the Youtube-VIS benchmark. Code is available at https://github.com/skrya/DefVIS.
Sudhir Yarram, Jialian Wu, Pan Ji, Yi Xu 0002, Junsong Yuan 0001
ICASSP2
2022 Multivehicle Object Tracking in Satellite Video Enhanced by Slow Features and Motion Features
abstract
With the development of video satellites, multimoving object tracking in satellite video is possible and has become a new challenging task. The difficulties are mainly caused by the characteristics of satellite videos: 1) small objects; 2) low contrast between objects and background; and 3) background in a state of continuous motion. These characteristics make it difficult for the advanced multiobject tracking algorithms in the natural video to give full play to their advantages, resulting in vast false alarms, missed objects, ID switches, and low-confidence bounding boxes. To tackle these problems, a novel multimoving object tracking method considering slow features (SFs) and motion features has been proposed in this research, named SF and motion feature-guided multiobject tracking (SFMFMOT), which realizes the continuous tracking of moving vehicles in satellite videos. A nonmaximum suppression (NMS) module guided by bounding box proposals based on SFs is designed to assist the object detection part by utilizing the sensitivity of SF analysis to the changed pixels. While removing a large number of static false alarms and supplementing missed objects, it improves the recall rate by increasing the confidence score of the correctly detected object bounding boxes. In order to improve the tracking performance, a set of optimization strategies based on motion features and time accumulation information are proposed to smooth the trajectory, remove static false alarms, and duplicate bounding boxes. The proposed method is evaluated in three satellite videos and its superiority is demonstrated.
Jialian Wu, Xin Su 0003, Qiangqiang Yuan, Huanfeng Shen, Liangpei Zhang 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 ForestDet: Large-Vocabulary Long-Tailed Object Detection and Instance Segmentation
abstract
Object detection and instance segmentation with a large number of object categories and long-tailed data distribution are challenging for most existing deep learning models. As the number of classes increases, the outputs of a classifier become sensitive to likely noisy logits, which can easily result in an incorrect recognition. To alleviate the large-vocabulary problem, we cluster fine-grained classes into coarser parent classes and then build a classification tree to classify an object into a fine-grained class via its parent class. Because the number of parent class is much fewer, their logits are more stable to suppress the wrong/noisy logits existed in the fine-grained class nodes. Due to a variety of ways for clustering fine-grained classes into parent classes, we can further construct multiple trees to build a classification forest where each single tree contributes its vote to the fine-grained classification. Moreover, a simple yet effective resampling method, termed as NMS Resampling, is proposed aiming at solving the long tail (data imbalance) problem. Our method, coined as ForestDet, serves as a plug-and-play module, which can be readily employed in both one-stage and two-stage object recognition models for recognizing more than 1000 categories. Extensive experiments are conducted on the large vocabulary dataset LVIS. Compared to the Mask R-CNN baseline, our two-stage counterpart Forest R-CNN significantly boosts the performance by 11.5% and 3.9% AP improvements on the rare categories and overall categories, respectively. Compared to the RetinaNet baseline, our one-stage counterpart Forest RetinaNet improves 2.1% AP on overall categories. Moreover, we achieve state-of-the-art results on the LVIS dataset.Code and models are available athttps://github.com/JialianW/Forest_RCNN.
Jialian Wu, Liangchen Song, Qian Zhang 0009, Ming Yang 0007, Junsong Yuan 0001
IEEE Trans. Multim.1
2021 Robust Knowledge Transfer via Hybrid Forward on the Teacher-Student Model
abstract
When adopting deep neural networks for a new vision task, a common practice is to start with fine-tuning some off-the-shelf well-trained network models from the community. Since a new task may require training a different network architecture with new domain data, taking advantage of off-the-shelf models is not trivial and generally requires considerable try-and-error and parameter tuning. In this paper, we denote a well-trained model as a teacher network and a model for the new task as a student network. We aim to ease the efforts of transferring knowledge from the teacher to the student network, robust to the gaps between their network architectures, domain data, and task definitions. Specifically, we propose a hybrid forward scheme in training the teacher-student models, alternately updating layer weights of the student model. The key merit of our hybrid forward scheme is on the dynamical balance between the knowledge transfer loss and task specific loss in training. We demonstrate the effectiveness of our method on a variety of tasks, e.g., model compression, segmentation, and detection, under a variety of knowledge transfer settings.
Liangchen Song, Jialian Wu, Ming Yang 0007, Qian Zhang 0009, Junsong Yuan 0001
AAAI2
2021 Track To Detect and Segment: An Online Multi-Object Tracker
abstract
Most online multi-object trackers perform object detection stand-alone in a neural net without any input from tracking. In this paper, we present a new online joint detection and tracking model, TraDeS (TRAck to DEtect and Segment), exploiting tracking clues to assist detection end-to-end. TraDeS infers object tracking offset by a cost volume, which is used to propagate previous object features for improving current object detection and segmentation. Effectiveness and superiority of TraDeS are shown on 4 datasets, including MOT (2D tracking), nuScenes (3D tracking), MOTS and Youtube-VIS (instance segmentation tracking). Project page: https://jialianwu.com/projects/TraDeS.html.
Jialian Wu, Jiale Cao, Liangchen Song, Yu Wang 0032, Ming Yang 0007, Junsong Yuan 0001
CVPR1
2021 Stacked Homography Transformations for Multi-View Pedestrian Detection
abstract
Multi-view pedestrian detection aims to predict a bird’s eye view (BEV) occupancy map from multiple camera views. This task is confronted with two challenges: how to establish the 3D correspondences from views to the BEV map and how to assemble occupancy information across views. In this paper, we propose a novel Stacked HOmography Transformations (SHOT) approach, which is motivated by approximating projections in 3D world coordinates via a stack of homographies. We first construct a stack of transformations for projecting views to the ground plane at different height levels. Then we design a soft selection module so that the network learns to predict the likelihood of the stack of transformations. Moreover, we provide an in-depth theoretical analysis on constructing SHOT and how well SHOT approximates projections in 3D world coordinates. SHOT is empirically verified to be capable of estimating accurate correspondences from individual views to the BEV map, leading to new state-of-the-art performance on standard evaluation benchmarks.
Liangchen Song, Jialian Wu, Ming Yang 0007, Qian Zhang 0009, Junsong Yuan 0001
ICCV2
2021 Handling Difficult Labels for Multi-label Image Classification via Uncertainty Distillation
abstract
Multi-label image classification aims to predict multiple labels for a single image. However, the difficulties of predicting different labels may vary dramatically due to semantic variations of the label as well as the image context. Direct learning of multi-label classification models has the risk of being biased and overfitting those difficult labels, e.g., deep network based classifiers are over-trained on the difficult labels, therefore, lead to false-positive errors of those difficult labels during testing. To handle difficult labels of multi-label image classification, we propose to calibrate the model, which not only predicts the labels but also estimates the uncertainty of the prediction. With the new calibration branch of the network, the classification model is trained with the pick-all-labels normalized loss and optimized pertaining to the number of positive labels. Moreover, to improve performance on difficult labels, instead of annotating them, we leverage the calibrated model as the teacher network and teach the student network about handling difficult labels via uncertainty distillation. Our proposed uncertainty distillation teaches the student network which labels are highly uncertain through prediction distribution distillation, and locates the image regions that cause such uncertain predictions through uncertainty attention distillation. Conducting extensive evaluations on benchmark datasets, we demonstrate that our proposed uncertainty distillation is valuable to handle difficult labels of multi-label image classification.
Liangchen Song, Jialian Wu, Ming Yang 0007, Qian Zhang 0009, Junsong Yuan 0001
ACM Multimedia2
2020 Temporal-Context Enhanced Detection of Heavily Occluded Pedestrians
abstract
State-of-the-art pedestrian detectors have performed promisingly on non-occluded pedestrians, yet they are still confronted by heavy occlusions. Although many previous works have attempted to alleviate the pedestrian occlusion issue, most of them rest on still images. In this paper, we exploit the local temporal context of pedestrians in videos and propose a tube feature aggregation network (TFAN) aiming at enhancing pedestrian detectors against severe occlusions. Specifically, for an occluded pedestrian in the current frame, we iteratively search for its relevant counterparts along temporal axis to form a tube. Then, features from the tube are aggregated according to an adaptive weight to enhance the feature representations of the occluded pedestrian. Furthermore, we devise a temporally discriminative embedding module (TDEM) and a part-based relation module (PRM), respectively, which adapts our approach to better handle tube drifting and heavy occlusions. Extensive experiments are conducted on three datasets, Caltech, NightOwls and KAIST, showing that our proposed method is significantly effective for heavily occluded pedestrian detection. Moreover, we achieve the state-of-the-art performance on the Caltech and NightOwls datasets.
Jialian Wu, Chunluan Zhou, Ming Yang 0007, Qian Zhang 0009, Junsong Yuan 0001
CVPR1
2020 Forest R-CNN: Large-Vocabulary Long-Tailed Object Detection and Instance Segmentation
abstract
Despite the previous success of object analysis, detecting and segmenting a large number of object categories with a long-tailed data distribution remains a challenging problem and is less investigated. For a large-vocabulary classifier, the chance of obtaining noisy logits is much higher, which can easily lead to a wrong recognition. In this paper, we exploit prior knowledge of the relations among object categories to cluster fine-grained classes into coarser parent classes, and construct a classification tree that is responsible for parsing an object instance into a fine-grained category via its parent class. In the classification tree, as the number of parent class nodes are significantly less, their logits are less noisy and can be utilized to suppress the wrong/noisy logits existed in the fine-grained class nodes. As the way to construct the parent class is not unique, we further build multiple trees to form a classification forest where each tree contributes its vote to the fine-grained classification. To alleviate the imbalanced learning caused by the long-tail phenomena, we propose a simple yet effective resampling method, NMS Resampling, to re-balance the data distribution. Our method, termed as Forest R-CNN, can serve as a plug-and-play module being applied to most object recognition models for recognizing more than 1000 categories. Extensive experiments are performed on the large vocabulary dataset LVIS. Compared with the Mask R-CNN baseline, the Forest R-CNN significantly boosts the performance with 11.5% and 3.9% AP improvements on the rare categories and overall categories, respectively. Moreover, we achieve state-of-the-art results on the LVIS dataset. Code is available at https://github.com/JialianW/Forest_RCNN.
Jialian Wu, Liangchen Song, Tiancai Wang, Qian Zhang 0009, Junsong Yuan 0001
ACM Multimedia1
2020 Self-Mimic Learning for Small-scale Pedestrian Detection
abstract
Detecting small-scale pedestrians is one of the most challenging problems in pedestrian detection. Due to the lack of visual details, the representations of small-scale pedestrians tend to be weak to be distinguished from background clutters. In this paper, we conduct an in-depth analysis of the small-scale pedestrian detection problem, which reveals that weak representations of small-scale pedestrians are the main cause for a classifier to miss them. To address this issue, we propose a novel Self-Mimic Learning (SML) method to improve the detection performance on small-scale pedestrians. We enhance the representations of small-scale pedestrians by mimicking the rich representations from large-scale pedestrians. Specifically, we design a mimic loss to force the feature representations of small-scale pedestrians to approach those of large-scale pedestrians. The proposed SML is a general component that can be readily incorporated into both one-stage and two-stage detectors, with no additional network layers and incurring no extra computational cost during inference. Extensive experiments on both the CityPersons and Caltech datasets show that the detector trained with the mimic loss is significantly effective for small-scale pedestrian detection and achieves state-of-the-art results on CityPersons and Caltech, respectively.
Jialian Wu, Chunluan Zhou, Qian Zhang 0009, Ming Yang 0007, Junsong Yuan 0001
ACM Multimedia1