EDBT 2026 Demo / reviewers in the wild / expert
Ruixin Liu
dblp:27/704
· DBLP profile ↗
25ranked-venue papers
6as first author
22since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 5 first-author · 14 since 2021Artificial intelligence and machine learning · 10 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Recognition of tea sprouts based on improved deeplabV3 + model
Ruixin Liu, Dengzhun Wang, Zhonghui Chen, Chunyu Yan, Jianwei Yan, Benliang Xie |
Multim. Tools Appl. | 1 |
| 2025 | Hierarchical Queries for 3D Lane Detection Based on Multi-Frame Point Cloudsabstract3D lane detection based on multi-frame point clouds is a critical task for autonomous driving. The challenge lies in efficiently performing temporal fusion using multiple data frames with incomplete yet complementary contexts. Existing methods either directly concatenate consecutive frames, avoiding intrinsic limitations of the raw data, or fuse entire feature maps, without distinguishing lane-related features from backgrounds. These solutions exhibit room for improvement in both precision and efficiency. In this paper, we propose an end-to-end lane detection network with hierarchical queries, which decodes lane features at different levels in a top-down manner for high-precision localization. This framework can be deployed on multi-frame inputs, as it efficiently achieves lane-related sequence fusion with reduced computational costs and improved inference speed. Specifically, we design semi-parametric lane geometry representations to model lanes as parametric curves and discrete points. Accordingly, hierarchical queries are proposed to focus on two-level lane geometries, including curve queries and point queries. Curve queries capture global structures of lanes projected onto the bird’s-eye-view (BEV) flat ground, while point queries aggregate multi-frame sequences obtained through curve-guided sampling, acquiring comprehensive and reliable point-level features. In the training stage, our proposed curve matching and point localization loss optimizes the detected lane geometries at both levels. Experiments conducted on the self-collected MultiBEV dataset validate that our method outperforms previously published single-frame and multi-frame methods. Codes are released at https://github.com/lrx02/HQNet Ruixin Liu, Zejian Yuan |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | Compact HD Map Construction via Douglas-Peucker Point TransformerabstractHigh-definition (HD) map construction requires a comprehensive understanding of traffic environments, encompassing centimeter-level localization and rich semantic information. Previous works face challenges in redundant point representation or high-complexity curve modeling. In this paper, we present a flexible yet effective map element detector that synthesizes hierarchical information with a compact Douglas-Peucker (DP) point representation in a transformer architecture for robust and reliable predictions. Specifically, our proposed representation approximates class-agnostic map elements with DP points, which are sparsely located in crucial positions of structures and can get rid of redundancy and complexity. Besides, we design a position constraint with uncertainty to avoid potential ambiguities. Moreover, pairwise-point shape matching constraints are proposed to balance local structural information of different scales. Experiments on the public nuScenes dataset demonstrate that our method overwhelms current SOTAs. Extensive ablation studies validate each component of our methods. Codes will be released at https://github.com/sweety121/DPFormer. Ruixin Liu, Zejian Yuan |
AAAI | 1 |
| 2024 | Offline handwritten mathematical expression recognition based on YOLOv5s
Hongbo Fang, Dengzhun Wang, Ruixin Liu, Qing Hou, Benliang Xie |
Vis. Comput. | 4 |
| 2023 | Flexible 3D Lane Detection by Hierarchical Shape MatchingabstractAs one of the basic while vital technologies for HD map construction, 3D lane detection is still an open problem due to varying visual conditions, complex typologies, and strict demands for precision. In this paper, an end-to-end flexible and hierarchical lane detector is proposed to precisely predict 3D lane lines from point clouds. Specifically, we design a hierarchical network predicting flexible representations of lane shapes at different levels, simultaneously collecting global instance semantics and avoiding local errors. In the global scope, we propose to regress parametric curves w.r.t adaptive axes that help to make more robust predictions towards complex scenes, while in the local vision the structure of lane segment is detected in each of the dynamic anchor cells sampled along the global predicted curves. Moreover, corresponding global and local shape matching losses and anchor cell generation strategies are designed. Experiments on two datasets show that we overwhelm current top methods under high precision standards, and full ablation studies also verify each part of our method. Our codes will be released at https://github.com/Doo-do/FHLD. Zhihao Guan, Ruixin Liu, Zejian Yuan, Ao Liu 0010, Erlong Li, Chao Zheng 0004, Shuqi Mei |
AAAI | 2 |
| 2023 | Learning to Detect 3D Lanes by Shape Matching and Embeddingabstract3D lane detection based on LiDAR point clouds is a challenging task that requires precise locations, accurate topologies, and distinguishable instances. In this paper, we propose a dual-level shape attention network (DSANet) with two branches for high-precision 3D lane predictions. Specifically, one branch predicts the refined lane segment shapes and the shape embeddings that encode the approximate lane instance shapes, the other branch detects the coarse-grained structures of the lane instances. In the training stage, two-level shape matching loss functions are introduced to jointly optimize the shape parameters of the twobranch outputs, which are simple yet effective for precision enhancement. Furthermore, a shape-guided segments aggregator is proposed to help local lane segments aggregate into complete lane instances, according to the differences of instance shapes predicted at different levels. Experiments conducted on our BEV-3DLanes dataset demonstrate that our method outperforms previous methods. Ruixin Liu, Zhihao Guan, Zejian Yuan, Ao Liu 0010, Tang Kun, Erlong Li, Chao Zheng 0004, Shuqi Mei |
WACV | 1 |
| 2023 | Separately Guided Context-Aware Network for Weakly Supervised Temporal Action Detection
Bairong Li, Yifan Pan, Ruixin Liu, Yuesheng Zhu |
Neural Process. Lett. | 3 |
| 2023 | Octave convolution-based vehicle detection using frame-difference as network inputabstractVehicle detection in video frames has been treated the same way detecting vehicle for an isolated image. However, the models designed for the isolated image are blind to fast-moving vehicles and cannot localize the moving targets partially occluded in the scene. In this case, we figure out a way to combine the classic moving target detection method with the neural network method. In this work, first, we propose to add three-differential-frames into the neural network of Yolov3 as the second input which contains the motion information on the front and back frames to detect vehicles partially occluded; second, we reform the network by using Octave Convolution to reduce memory and computational cost while boosting accuracy. We experimentally show that by using the aforementioned methods together, compared with using original YOLOv3 on UA-DETRAC data set, AP is increased by 2.31%, recall is increased by 4.01%, and precision is increased by 3.10%. We demonstrate that the proposed method is indeed effective. Ruixin Liu, Zhonghui Chen, Dengzhun Wang, Benliang Xie |
Vis. Comput. | 2 |
| 2022 | Interactive Image Inpainting Using Semantic GuidanceabstractImage inpainting approaches have achieved significant progress with the help of deep neural networks. How-ever, existing approaches mainly focus on leveraging the priori distribution learned by neural networks to produce a single inpainting result or further yielding multiple solutions, where the controllability is not well studied. This paper develops a novel image inpainting approach that enables users to customize the inpainting result by their own preference or memory. Specifically, our approach is composed of two stages that utilize the prior of neural network and user’s guidance to jointly inpaint corrupted images. In the first stage, an autoencoder based on a novel external spatial attention mechanism is deployed to produce reconstructed features of the corrupted image and a coarse inpainting result that provides semantic mask as the medium for user interaction. In the second stage, a semantic decoder that takes the reconstructed features as prior is adopted to synthesize a fine inpainting result guided by user’s customized semantic mask, so that the final inpainting result will share the same content with user’s guidance while the textures and colors reconstructed in the first stage are preserved. Extensive experiments demonstrate the superiority of our approach in terms of inpainting quality and controllability. Wangbo Yu, Jinhao Du, Ruixin Liu, Yuesheng Zhu |
ICPR | 3 |
| 2022 | Unsupervised Domain Adaptation Integrating Transformer and Mutual Information for Cross-Corpus Speech Emotion RecognitionabstractThis paper focuses on an interesting task, i.e., unsupervised cross-corpus Speech Emotion Recognition (SER), in which the labelled training (source) corpus and the unlabelled testing (target) corpus have different feature distributions, resulting in the discrepancy between the source and target domains. To address this issue, this paper proposes an unsupervised domain adaptation method integrating Transformers and Mutual Information (MI) for cross-corpus SER. Initially, our method employs encoder layers of Transformers to capture long-term temporal dynamics in an utterance from the extracted segment-level log-Mel spectrogram features, thereby producing the corresponding utterance-level features for each utterance in two domains. Then, we propose an unsupervised feature decomposition method with a hybrid Max-Min MI strategy to separately learn domain-invariant features and domain-specific features from the extracted mixed utterance-level features, in which the discrepancy between two domains is eliminated as much as possible and meanwhile their individual characteristic is preserved. Finally, an interactive Multi-Head attention fusion strategy is designed to learn the complementarity between domain-invariant features and domain-specific features so that they can be interactively fused for SER. Extensive experiments on the IEMOCAP and MSP-Improv datasets demonstrate the effectiveness of our proposed method on unsupervised cross-corpus SER tasks, outperforming state-of-the-art unsupervised cross-corpus SER methods. Shiqing Zhang, Ruixin Liu, Yijiao Yang, Xiaoming Zhao 0002, Jun Yu 0002 |
ACM Multimedia | 2 |
| 2022 | Weakly Supervised Temporal Action Detection With Temporal Dependency LearningabstractWeakly supervised temporal action detection aims at localizing temporal positions of action instances in untrimmed videos with only action class labels. In general, previous methods individually classify each frame based on the appearance information and the short-term motion information, and then integrate consecutive high-response action frames into entities which serve as detected action instances. However, the long-range temporal dependencies between action frames are not fully utilized, and the detection results are more likely to be trapped in the most discriminative action segments. To alleviate this issue, we propose a novel two-branch (i.e., the coarse detection branch and the refining detection branch) detection framework with learning the long-range temporal dependencies for obtaining more accurate detection results, where only action class labels are required. The coarse detection branch is used to localize the most discriminative segments of action instances based on a typical multi-instance learning paradigm under the supervision of action class labels, whereas the refining detection branch is expected to localize the less discriminative segments of action instances via learning the long-range temporal dependencies between frames based on the proposed Transformer-style architecture and learning strategies. This collaboration mechanism takes full advantage of complementary information from the provided action class labels and the natural temporal dependencies between action frames, forming a more comprehensive solution. Consequently, our method obtains more precise detection results. Expectedly, the proposed method outperforms recent weakly supervised temporal action detection methods on dataset THUMOS14 and ActivityNet measured by mAP@tIoU and AR@AN. Bairong Li, Ruixin Liu, Tianquan Chen, Yuesheng Zhu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Temporal Group Fusion Network for Deep Video InpaintingabstractVideo inpainting is a task of synthesizing spatio-temporal coherent content in missing regions of the given video sequence, which has recently drawn increasing attention. To utilize the temporal information across frames, most recent deep learning-based methods align reference frames to target frame firstly with explicit or implicit motion estimation and then integrate the information from the aligned frames. However, their performance relies heavily on the accuracy of frame-to-frame alignment. To alleviate the above problem, in this paper, a novel Temporal Group Fusion Network (TGF-Net) is proposed to effectively integrate temporal information through a two-stage fusion strategy. Specifically, the input frames are reorganized into different groups, where each group is followed by an intra-group fusion module to integrate information within the group. Different groups provide complementary information for the missing region. A temporal attention model is further designed to adaptively integrate the information across groups. Such a temporal information fusion way gets rid of the dependence on alignment operations, greatly improving the visual quality and temporal consistency of the inpainted results. In addition, a coarse alignment model is introduced at the beginning of the network to handle videos with large motion. Extensive experiments on DAVIS and Youtube-VOS datasets demonstrate the superiority of our proposed method in terms of PSNR/SSIM values, visual quality and temporal consistency, respectively. Ruixin Liu, Bairong Li, Yuesheng Zhu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Semantic-Aware Context Aggregation for Image InpaintingabstractRecent attention-based image inpainting methods have made inspiring progress by propagating distant contextual information into holes. However, they tend to generate blurry contents since the propagation process is always misled by preliminarily-recovered holes features which are not well-inferred. To handle this problem, we propose a novel semantic-aware context aggregation module (SACA) that aggregates distant contextual information from a semantic perspective by exploiting the internal semantic similarity of the input feature map. Compared with existing attention mechanisms that model the relation of all pixel-pairs, SACA can suppress the impact of misleading holes features in context aggregation and significantly reduce computation burden by learning the relation between pixels and semantics. Also, we apply SACA to both high-level and low-level feature maps in our model for generating both semantically and visually plausible results. Extensive experiments on Outdoor Scenes, CelebA and Paris StreetView datasets validate the superiority of our method compared with existing methods. Zhilin Huang, Chujun Qin, Ruixin Liu, Zhenyu Weng, Yuesheng Zhu |
ICASSP | 3 |
| 2021 | SERN: Stance Extraction and Reasoning Network for Fake News DetectionabstractFake news brings us panic and misunderstanding against the truth, especially under some unusual circumstances, such as the outbreak of COVID-19. It’s crucial to detect fake news on social media early to avoid further propagation. Previous methods manually label the stances implied in post-reply pairs to aid fake news detection, which costs much time and effort. To solve this problem, a novel Stance Extraction and Reasoning Network (SERN) is proposed to extract the stances implied in post-reply pairs implicitly and integrate the stance representations for fake news detection without manually labeling stances, which saves much time and effort. Besides, the adequate utilization of multimodal content in the news is beneficial for complementing information for unimodal representation and jointly improving decision confidence. Thus, a sentence-guided visual attention mechanism is proposed in the text-image fusion module that leverages text-image content for better fake news detection. Encouraging empirical results on Fakeddit and PHEME demonstrate that our method outperforms the state-of-the-art methods. Jianhui Xie, Ruixin Liu, Yinghong Zhang, Yuesheng Zhu |
ICASSP | 3 |
| 2021 | GI-AEE: GAN Inversion Based Attentive Expression Embedding Network For Facial Expression EditingabstractFacial expression editing aims to modify facial expression by specific conditions. Existing methods adopt an encoder-decoder architecture under the guidance of expression condition to process the desired expression. However, these methods always tend to produce artifacts and blurs in expression-intensive regions due to simultaneously modifying images in expression changed regions and ensuring the consistency of other attributes with the source image. To address these issues, we propose a GAN inversion based Attentive Expression Embedding Network (GI-AEE) for facial expression editing, which decouples this task utilizing GAN inversion to alleviate the strong effect of the source image on the target image and produces high-quality expression editing results. Furthermore, different from existing methods that directly embed the expression condition into the network, we propose an Attentive Expression Embedding module to embed corresponding expression vectors into different facial regions, producing more plausible results. Qualitative and quantitative experiments demonstrate our method outperforms the state-of-the-art expression editing methods. Ruixin Liu, Yifan Pan, Yuesheng Zhu, Zhiqiang Bai |
ICIP | 2 |
| 2021 | Manipulation-Invariant Fingerprints for Cross-Dataset Deepfake Detection
Zuoyan Li, Ruixin Liu, Yuesheng Zhu |
ICONIP (4) | 3 |
| 2021 | Bi-encoder Network with Structure-texture Consistency for Image InpaintingabstractExisting image inpainting methods have shown their potential in filling corrupted regions with plausible contents. However, these methods tend to produce results with distorted structures or unnatural textures since they neglect the difference between structures and textures in images and jointly process these two different types of information. To solve this problem, we propose a bi-encoder network (BE-Net) that seeks to handle structure and texture information separately, and fuse them to reconstruct completed images. Specifically, BE-Net first uses two parallel encoders to infer structure and texture features of the input images respectively. Then a structure-texture consistency module (STCM) is designed to weaken artifacts and enhance visual coherency of the output images by keeping the texture features consistent with the structure features. Finally, the structure features and the texture features are fused at each level of the decoder to recover images with reasonable structures and realistic textures. Extensive experiments on Paris StreetView and CelebA datasets show the proposed approach is effective in generating realistic and visually plausible results and outperforms several state-of-the-art methods. Chujun Qin, Zhilin Huang, Ruixin Liu, Zhenyu Weng, Yuesheng Zhu |
IJCNN | 3 |
| 2021 | Confidence-Based Global Attention Guided Network for Image Inpainting
Zhilin Huang, Chujun Qin, Ruixin Liu, Yuesheng Zhu |
MMM (1) | 4 |
| 2021 | One-sided Versus Two-sided: A Novel Opinion Dynamics Information-Type Education-Based Hegselmann-Krause ModelabstractThe classic Hegselmann–Krause opinion dynamics model plays an important role in analyzing opinion evolution among people. However, it is not in line with the reality of interpersonal communication during the modern online social times, in which people are easy to exchange opinions with another one who agrees or disagrees with themselves on certain issues by social networks. In this paper, considering the effects of the one-sided versus two-sided presentation of a controversial issue, we improve the classic HK model and propose a novel information-type education-based Hegselmann–Krause model. In our work, we formulate rules of how agents’ opinions change and carry out simulation experiments on networks with different proportions of high-educated agents. Extensive experiments on artificially generated networks show that it is more difficult for someone to reach a consensus with another one who is within a group with a higher proportion of low-educated people than with a higher proportion of high-educated people, which verify the effectiveness of the proposed model. Minghua Xu 0001, Ziling Luo, Ruixin Liu, Bang Wang 0001, Han Xu 0003 |
SMC | 3 |
| 2021 | COVID-19 Vaccine Sensing: Sentiment Analysis from Twitter DataabstractThe COVID-19 outbreak a pandemic, which poses a serious threat to global public health and lead to a tsunami of online social media. Individuals frequently express their views, opinions and emotions about the events of the pandemic on Twitter, Facebook, etc. Many researches try to analyze the sentiment of the COVID-19-related content from these social networks. However, they have rarely focused on the vaccine. In this paper, we study the COVID-19 vaccine topic from Twitter. Specifically, all the tweets related to COVID-19 vaccine from December 15th, 2020 to February 10th, 2021 are collected by using the Twitter API, then the unsupervised learning VADER model is used to judge the emotion categories (positive, neutral, negative) and calculate the sentiment value of the dataset. Based on the interaction between users, a communication topological network is constructed and the emotional direction is explored. We find that people had different sentiments between Chinese vaccine and those in other countries. The sentiment value might be affected by the number of daily news cases and deaths, the nature of key issues in the communication network. And revealing that the key nodes in the social network can produce emotional contagion to other nodes. Han Xu 0003, Ruixin Liu, Ziling Luo, Minghua Xu 0001, Bang Wang 0001 |
SMC | 2 |
| 2021 | Learning frame-level affinity with video-level labels for weakly supervised temporal action detection
Bairong Li, Yuesheng Zhu, Ruixin Liu, Zhenyu Weng |
Neurocomputing | 3 |
| 2021 | Bidirectional loss function for Label Enhancement and distribution learning
Xinyuan Liu 0001, Jihua Zhu, Qinghai Zheng, Zhongyu Li 0002, Ruixin Liu, Jun Wang 0024 |
Knowl. Based Syst. | 5 |
| 2020 | Temporal Adaptive Alignment Network for Deep Video InpaintingabstractVideo inpainting aims to synthesize visually pleasant and temporally consistent content in missing regions of video. Due to a variety of motions across different frames, it is highly challenging to utilize effective temporal information to recover videos. Existing deep learning based methods usually estimate optical flow to align frames and thereby exploit useful information between frames. However, these methods tend to generate artifacts once the estimated optical flow is inaccurate. To alleviate above problem, we propose a novel end-to-end Temporal Adaptive Alignment Network(TAAN) for video inpainting. The TAAN aligns reference frames with target frame via implicit motion estimation at a feature level and then reconstruct target frame by taking the aggregated aligned reference frame features as input. In the proposed network, a Temporal Adaptive Alignment (TAA) module based on deformable convolutions is designed to perform temporal alignment in a local, dense and adaptive manner. Both quantitative and qualitative evaluation results show that our method significantly outperforms existing deep learning based methods. Ruixin Liu, Zhenyu Weng, Yuesheng Zhu, Bairong Li |
IJCAI | 1 |
| 2018 | A training samples selection method based on system identification for STAP
Huiyong Li 0001, Weiwei Bao, Jinfeng Hu, Julan Xie, Ruixin Liu |
Signal Process. | 5 |
| 2006 | The Minimum Detectable Capacity of Digital Image Information Hiding
Fan Zhang 0028, Ruixin Liu, Xinhong Zhang |
ISNN (2) | 2 |