VLDB 2026 Research / reviewers in the wild / expert
Ruijin Liu
dblp:254/7956
· DBLP profile ↗
9ranked-venue papers
5as first author
8since 2021 · last 2024
0000-0002-3672-5332ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Align Before Adapt: Leveraging Entity-to-Region Alignments for Generalizable Video Action RecognitionabstractLarge-scale visual-language pre-trained models have achieved significant success in various video tasks. However, most existing methods follow an “adapt then align” paradigm, which adapts pre-trained image encoders to model video-level representations and utilizes one-hot or text embedding of the action labels for supervision. This paradigm overlooks the challenge of mapping from static images to complicated activity concepts. In this paper, we propose a novel “Align before Adapt” (ALT) paradigm. Prior to adapting to video representation learning, we exploit the entity-to-region alignments for each frame. The alignments are fulfilled by matching the region-aware image embeddings to an offline-constructed text corpus. With the aligned entities, we feed their text embeddings to a transformer-based video adapter as the queries, which can help extract the semantics of the most important entities from a video to a vector. This paradigm reuses the visual-language alignment of VLP during adaptation and tries to explain an action by the underlying entities. This helps understand actions by bridging the gap with complex activity semantics, particularly when facing unfamiliar or unseen categories. ALT demonstrates competitive performance while maintaining remarkably low computational costs. In fully supervised experiments, it achieves 88.1 % top-1 accuracy on Kinetics-400 with only 4947 GFLOPs. Moreover, ALT outperforms the previous state-of-the-art methods in both zero-shot and fewshot experiments, emphasizing its superior generalizability across various learning scenarios. Yifei Chen 0010, Dapeng Chen, Ruijin Liu, Wenyuan Xue, Wei Peng 0011 |
CVPR | 3 |
| 2023 | Video Action Recognition with Attentive Semantic UnitsabstractVisual-Language Models (VLMs) have significantly advanced video action recognition. Supervised by the semantics of action labels, recent works adapt the visual branch of VLMs to learn video representations. Despite the effectiveness proved by these works, we believe that the potential of VLMs has yet to be fully harnessed. In light of this, we exploit the semantic units (SU) hiding behind the action labels and leverage their correlations with fine-grained items in frames for more accurate action recognition. SUs are entities extracted from the language descriptions of the entire action set, including body parts, objects, scenes, and motions. To further enhance the alignments between visual contents and the SUs, we introduce a multi-region attention module (MRA) to the visual branch of the VLM. The MRA allows the perception of region-aware visual features beyond the original global feature. Our method adaptively attends to and selects relevant SUs with visual features of frames. With a cross-modal decoder, the selected SUs serve to decode spatiotemporal video representations. In summary, the SUs as the medium can boost discriminative ability and transferability. Specifically, in fully-supervised learning, our method achieved 87.8% top-1 accuracy on Kinetics-400. In K=2 few-shot experiments, our method surpassed the previous state-of-the-art by +7.1% and +15.0% on HMDB-51 and UCF-101, respectively. Yifei Chen 0010, Dapeng Chen, Ruijin Liu, Wei Peng 0011 |
ICCV | 3 |
| 2023 | PBFormer: Capturing Complex Scene Text Shape with Polynomial Band TransformerabstractWe present PBFormer, an efficient yet powerful scene text detector that unifies the transformer with a novel text shape representation Polynomial Band (PB). The representation has four polynomial curves to fit a text's top, bottom, left, and right sides, which can capture a text with a complex shape by varying polynomial coefficients. PB has appealing features compared with conventional representations: 1) It can model different curvatures with a fixed number of parameters, while polygon-points-based methods need to utilize a different number of points. 2) It can distinguish adjacent or overlapping texts as they have apparent different curve coefficients, while segmentation-based or points-based methods suffer from adhesive spatial positions. PBFormer combines the PB with the transformer, which can directly generate smooth text contours sampled from predicted curves without interpolation. A parameter-free cross-scale pixel attention (CPA) module is employed to highlight the feature map of a suitable scale while suppressing the other feature maps. The simple operation can help detect small-scale texts and is compatible with the one-stage DETR framework, where no postprocessing exists for NMS. Furthermore, PBFormer is trained with a shape-contained loss, which not only enforces the piecewise alignment between the ground truth and the predicted curves but also makes curves' position and shapes consistent with each other. Without bells and whistles about text pre-training, our method is superior to the previous state-of-the-art text detectors on the arbitrary-shaped text datasets. Codes will be public. Ruijin Liu, Ning Lu 0003, Dapeng Chen, Cheng Li 0040, Zejian Yuan, Wei Peng 0011 |
ACM Multimedia | 1 |
| 2022 | Learning to Predict 3D Lane Shape and Camera Pose from a Single Image via Geometry ConstraintsabstractDetecting 3D lanes from the camera is a rising problem for autonomous vehicles. In this task, the correct camera pose is the key to generating accurate lanes, which can transform an image from perspective-view to the top-view. With this transformation, we can get rid of the perspective effects so that 3D lanes would look similar and can accurately be fitted by low-order polynomials. However, mainstream 3D lane detectors rely on perfect camera poses provided by other sensors, which is expensive and encounters multi-sensor calibration issues. To overcome this problem, we propose to predict 3D lanes by estimating camera pose from a single image with a two-stage framework. The first stage aims at the camera pose task from perspective-view images. To improve pose estimation, we introduce an auxiliary 3D lane task and geometry constraints to benefit from multi-task learning, which enhances consistencies between 3D and 2D, as well as compatibility in the above two tasks. The second stage targets the 3D lane task. It uses previously estimated pose to generate top-view images containing distance-invariant lane appearances for predicting accurate 3D lanes. Experiments demonstrate that, without ground truth camera pose, our method outperforms the state-of-the-art perfect-camera-pose-based methods and has the fewest parameters and computations. Codes are available at https://github.com/liuruijin17/CLGo. Ruijin Liu, Dapeng Chen, Zhiliang Xiong, Zejian Yuan |
AAAI | 1 |
| 2022 | Learning TBox With a Cascaded Anchor-Free Network for Vehicle DetectionabstractVehicle detection, the process of identifying vehicles as axis-aligned bounding boxes in still images, is widely used to estimate the range, time-to-collision, and motion of autonomous vehicles (AVs). Bounding boxes, while convenient, are too coarse to adapt well to vehicle shape and pose variations. In this work, we presentTBox(Trapezoid & Box), a novel fine-grained representation useful for both localization and recognition that extends the bounding box by restricting the spatial extent of a vehicle to a set of keypoints and indicating semantically significant local areas using subclasses. In contrast to the previous monolithic models, we propose a cascaded anchor-free architecture to estimate the bounding box and TBox. One subnetwork uses a stacked hourglass network to detect each vehicle as a pair of corners without using anchors. Specifically, it learns corner affinity fields, enabling it to perform robust corner grouping. The other subnetwork estimates a TBox as a set of keypoints. This subnetwork utilizes the bounding box results to avoid ambiguous keypoint associations and reuses existing features to reduce the number of parameters. We also propose a multitask learning strategy for training the cascaded model that implicitly integrates the global context with local details, introducing improvements for both tasks. During testing, a refinement algorithm explicitly uses robust local keypoints to correct possible global box errors, ensuring tight geometric representations for nearby critical vehicles. The experiments show that our method outperforms existing anchor-free detectors for vehicle detection and achieves better performance on the TBox task while using a small model. Ruijin Liu, Zejian Yuan |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2021 | Order-independent Matching with Shape Similarity for Parking Slot Detection
Ziyi Yin 0001, Ruijin Liu, Zejian Yuan, Zhiliang Xiong |
BMVC | 2 |
| 2021 | Multimodal Transformer Networks for Pedestrian Trajectory PredictionabstractWe consider the problem of forecasting the future locations of pedestrians in an ego-centric view of a moving vehicle. Current CNNs or RNNs are flawed in capturing the high dynamics of motion between pedestrians and the ego-vehicle, and suffer from the massive parameter usages due to the inefficiency of learning long-term temporal dependencies. To address these issues, we propose an efficient multimodal transformer network that aggregates the trajectory and ego-vehicle speed variations at a coarse granularity and interacts with the optical flow in a fine-grained level to fill the vacancy of highly dynamic motion. Specifically, a coarse-grained fusion stage fuses the information between trajectory and ego-vehicle speed modalities to capture the general temporal consistency. Meanwhile, a fine-grained fusion stage merges the optical flow in the center area and pedestrian area, which compensates the highly dynamic motion of ego-vehicle and target pedestrian. Besides, the whole network is only attention-based that can efficiently model long-term sequences for better capturing the temporal variations. Our multimodal transformer is validated on the PIE and JAAD datasets and achieves state-of-the-art performance with the most light-weight model size. The codes are available at https://github.com/ericyinyzy/MTN_trajectory. Ziyi Yin 0001, Ruijin Liu, Zhiliang Xiong, Zejian Yuan |
IJCAI | 2 |
| 2021 | End-to-end Lane Shape Prediction with TransformersabstractLane detection, the process of identifying lane markings as approximated curves, is widely used for lane departure warning and adaptive cruise control in autonomous vehicles. The popular pipeline that solves it in two steps- feature extraction plus post-processing, while useful, is too inefficient and flawed in learning the global context and lanes' long and thin structures. To tackle these issues, we propose an end-to-end method that directly outputs parameters of a lane shape model, using a network built with a transformer to learn richer structures and context. The lane shape model is formulated based on road structures and camera pose, providing physical interpretation for parameters of network output. The transformer models non-local interactions with a self-attention mechanism to capture slender structures and global context. The proposed method is validated on the TuSimple benchmark and shows state-of-the-art accuracy with the most lightweight model size and fastest speed. Additionally, our method shows excellent adaptability to a challenging self-collected lane detection dataset, showing its powerful deployment potential in real applications. Codes are available at https://github.com/liuruijin17/LSTR. Ruijin Liu, Zejian Yuan, Zhiliang Xiong |
WACV | 1 |
| 2019 | Beyond Bounding Box: Fine-Grained Vehicle Detection via Single Stage Detector with Hierarchical outputabstractVehicle detection is a crucial module of the camera-based forward collision alert, which is usually used to calculate range and time-to-collision. Conventional methods based on bounding box are too coarse to handle challenging situations such as vehicle pose variations. In this paper, we propose a novel vehicle detector with a fine-grained output representation. The detector exploits a hierarchical tree-like output representation by introducing two subclasses and virtual control points, which not only discriminates each face of a vehicle but also locates their boundaries accurately. Our detector adopts popular single stage multi-scale CNN framework, which is equipped with the hierarchical output, and is beyond the bounding box methods. Experiments on our large-scale self-collected dataset show that our method achieves satisfactory performance. Ruijin Liu, Fangying Luo, Zejian Yuan |
ICIP | 1 |