Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Xuejian Rong

dblp:150/6735 · DBLP profile ↗
← Back
17ranked-venue papers
6as first author
5since 2021 · last 2022
0000-0001-6617-9582ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 10 · 4 first-author · 4 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
4 papers
Image and video processing · 40% Multimedia analysis and retrieval · 33% Rendering · 27%
Artificial intelligence
5 papers
3D vision · 73% Vision and language · 13% Segmentation and scene understanding · 11%
Human-computer interaction and pervasive computing
2 papers
Accessibility and assistive technology · 96% Ubiquitous computing and smart environments · 4%

Topics — the 21 heaviest of 25, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Multimedia analysis and retrieval › image retrieval
scene text retrieval
0.922022
Unambiguous Text Localization, Retrieval, and Recognition for Cluttered Scenes · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Unambiguous Text Localization and Retrieval for Cluttered Scenes · CVPR 2017
Accessibility and assistive technology
assistive navigation
0.622019
Vision-Based Mobile Indoor Assistive Navigation Aid for Blind People · IEEE Trans. Mob. Comput. 2019
Demo: Assisting Visually Impaired People Navigate Indoors · IJCAI 2016
Accessibility and assistive technology › assistive navigation
indoor navigation for visual impairments
0.622019
Vision-Based Mobile Indoor Assistive Navigation Aid for Blind People · IEEE Trans. Mob. Comput. 2019
Demo: Assisting Visually Impaired People Navigate Indoors · IJCAI 2016
Rendering
neural radiance fields
0.612022
Boosting View Synthesis with Residual Transfer · CVPR 2022
Rendering
novel view synthesis
0.612022
Boosting View Synthesis with Residual Transfer · CVPR 2022
Image and video processing › document image analysis › text detection and recognition
scene text recognition
0.612022
Unambiguous Text Localization, Retrieval, and Recognition for Cluttered Scenes · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Computer vision › 3D vision
camera pose estimation
0.512021
Robust Consistent Video Depth Estimation · CVPR 2021
Computer vision › 3D vision
depth estimation
0.512021
Robust Consistent Video Depth Estimation · CVPR 2021
Computer vision › 3D vision › depth estimation › video depth estimation
monocular video depth estimation
0.512021
Robust Consistent Video Depth Estimation · CVPR 2021
Computer vision › Vision and language › visual grounding
referring expression comprehension
0.412020
Unambiguous Scene Text Segmentation With Referring Expression Comprehension · IEEE Trans. Image Process. 2020
Computer vision › Segmentation and scene understanding › image segmentation
scene text segmentation
0.412020
Unambiguous Scene Text Segmentation With Referring Expression Comprehension · IEEE Trans. Image Process. 2020
Image and video processing › image restoration › image denoising › camera noise removal
burst denoising
0.412020
Burst Denoising via Temporally Shifted Wavelet Transforms · ECCV (13) 2020
Image and video processing › image restoration
image denoising
0.412020
Burst Denoising via Temporally Shifted Wavelet Transforms · ECCV (13) 2020
Computer vision › 3D vision
3d scene reconstruction
0.412019
Incremental Scene Synthesis · NeurIPS 2019
Computer vision › 3D vision
3d scene understanding
0.412019
Incremental Scene Synthesis · NeurIPS 2019
Computer vision › 3D vision › 3d scene understanding
scene synthesis
0.412019
Incremental Scene Synthesis · NeurIPS 2019
Image and video processing › pattern detection
scene text detection
0.312017
Unambiguous Text Localization and Retrieval for Cluttered Scenes · CVPR 2017
Computer vision › 3D vision
3d reconstruction
0.112021
Robust Consistent Video Depth Estimation · CVPR 2021
Computer vision › 3D vision › visual localization
semantic localization
0.112019
Vision-Based Mobile Indoor Assistive Navigation Aid for Blind People · IEEE Trans. Mob. Comput. 2019
Robotics › Robot navigation and mapping
SLAM
0.112019
Incremental Scene Synthesis · NeurIPS 2019
Ubiquitous computing and smart environments
indoor localization
0.112016
Demo: Assisting Visually Impaired People Navigate Indoors · IJCAI 2016

Methods — techniques the papers use, named apart from their topics

context reasoning · 0.9residual transfer · 0.6recurrent neural network · 0.6heuristic weighting · 0.6convolutional representations · 0.6geometry-aware depth filtering · 0.5geometric optimization · 0.5deformation splines · 0.5convolutional neural network · 0.5wavelet transform · 0.4visual-linguistic joint modeling · 0.4saliency response map · 0.4deep network · 0.4visual positioning service · 0.4sampling · 0.4learned scene prior · 0.4kalman filter · 0.4hallucination · 0.4
YearPublicationVenuePosition
2022 Boosting View Synthesis with Residual Transfer
abstract
Volumetric view synthesis methods with neural representations, such as NeRF and NeX, have recently demonstrated high-quality novel view synthesis. However, optimizing these representations is slow, and even fully trained models cannot reproduce all fine details in the input views. We present a simple but effective technique to boost the rendering quality, which can be easily integrated with most view synthesis methods. The core idea is to transfer color resid-uals (the difference between the input images and their re-construction) from training views to novel views. We blend the residuals from multiple views using a heuristic weighting scheme depending on ray visibility and angular differ-ences. We integrate our technique with several state-of-the-art view synthesis methods and evaluate the Real Forward-facing and the Shiny datasets. Our results show that at about 1/10th the number of training iterations, we achieve the same rendering quality as fully converged NeRF and NeX models, and when applied to fully converged models, we significantly improve their rendering quality.
Xuejian Rong, Jia-Bin Huang 0001, Ayush Saraf, Changil Kim 0001, Johannes Kopf 0001
CVPR1
2022 Unambiguous Text Localization, Retrieval, and Recognition for Cluttered Scenes
abstract
Text instance as one category of self-described objects provides valuable information for understanding and describing cluttered scenes. The rich and precise high-level semantics embodied in the text could drastically benefit the understanding of the world around us. While most recent visual phrase grounding approaches focus on general objects, this paper explores extracting designated texts and predicting unambiguous scene text information, i.e., to accurately localize and recognize a specific targeted text instance in a cluttered image from natural language descriptions (referring expressions). To address this issue, first a novel recurrent dense text localization network (DTLN) is proposed to sequentially decode the intermediate convolutional representations of a cluttered scene image into a set of distinct text instance detections. Our approach avoids repeated text detections at multiple scales by recurrently memorizing previous detections, and effectively tackles crowded text instances in close proximity. Second, we propose a context reasoning text retrieval (CRTR) model, which jointly encodes text instances and their context information through a recurrent network, and ranks localized text bounding boxes by a scoring function of context compatibility. Third, a recurrent text recognition module is introduced to extend the applicability of aforementioned DTLN and CRTR models, via text verification or transcription. Quantitative evaluations on standard scene text extraction benchmarks and a newly collected scene text retrieval dataset demonstrate the effectiveness and advantages of our models for the joint scene text localization, retrieval, and recognition task.
Xuejian Rong, Chucai Yi, Yingli Tian
IEEE Trans. Pattern Anal. Mach. Intell.1
2021 AMICO: Amodal Instance Composition
Peiye Zhuang, Denis Demandolx, Ayush Saraf, Xuejian Rong, Changil Kim 0001, Jia-Bin Huang 0001
BMVC4
2021 Robust Consistent Video Depth Estimation
abstract
We present an algorithm for estimating consistent dense depth maps and camera poses from a monocular video. We integrate a learning-based depth prior, in the form of a convolutional neural network trained for single-image depth estimation, with geometric optimization, to estimate a smooth camera trajectory as well as detailed and stable depth reconstruction. Our algorithm combines two complementary techniques: (1) flexible deformation-splines for low-frequency large-scale alignment and (2) geometry-aware depth filtering for high-frequency alignment of fine depth details. In contrast to prior approaches, our method does not require camera poses as input and achieves robust reconstruction for challenging hand-held cell phone captures containing a significant amount of noise, shake, motion blur, and rolling shutter deformations. Our method quantitatively outperforms state-of-the-arts on the Sintel benchmark for both depth and pose estimations and attains favorable qualitative results across diverse wild datasets.
Johannes Kopf 0001, Xuejian Rong, Jia-Bin Huang 0001
CVPR2
2021 Self-supervised 4D Spatio-temporal Feature Learning via Order Prediction of Sequential Point Cloud Clips
abstract
Recently 3D scene understanding attracts attention for many applications, however, annotating a vast amount of 3D data for training is usually expensive and time consuming. To alleviate the needs of ground truth, we propose a self-supervised schema to learn 4D spatio-temporal features (i.e. 3 spatial dimensions plus 1 temporal dimension) from dynamic point cloud data by predicting the temporal order of sampled and shuffled point cloud clips. 3D sequential point cloud contains precious geometric and depth information to better recognize activities in 3D space compared to videos. To learn the 4D spatio-temporal features, we introduce 4D convolution neural networks to predict the temporal order on a self-created large scale dataset, NTU-PCLs, derived from the NTU-RGB+D dataset. The efficacy of the learned 4D spatio-temporal features is verified on two tasks: 1) Self-supervised 3D nearest neighbor retrieval; and 2) Self-supervised representation learning transferred for action recognition on a smaller 3D dataset. Our extensive experiments prove the effectiveness of the proposed self-supervised learning method which achieves comparable results w.r.t. the fully-supervised methods on action recognition on MSRAction3D dataset.
Haiyan Wang 0019, Xuejian Rong, Jinglun Feng, Yingli Tian
WACV3
2020 Burst Denoising via Temporally Shifted Wavelet Transforms
Xuejian Rong, Denis Demandolx, Kevin Matzen, Priyam Chatterjee, Yingli Tian
ECCV (13)1
2020 Towards Efficient 3D Point Cloud Scene Completion via Novel Depth View Synthesis
abstract
3D point cloud completion has been a long-standing challenge at scale, and corresponding per-point supervised training strategies suffered from cumbersome annotations. 2D supervision has recently emerged as a promising alternative for 3D tasks, but specific approaches for 3D point cloud completion still remain to be explored. To overcome these limitations, we propose an end-to-end method that directly lifts a single depth map to a completed point cloud. With one depth map as input, a multiway novel depth view synthesis network (NDVNet) is designed to infer coarsely completed depth maps under various viewpoints. Meanwhile, a geometric depth perspective rendering module is introduced to utilize the raw input depth map to generate a re-projected depth map for each view. Therefore, the two parallelly generated depth maps for each view are further concatenated and refined by a depth completion network (DCNet). The final completed point cloud is fused from all refined depth views. Experimental results demonstrate the effectiveness of our proposed approach composed of aforementioned components, to produce high-quality, state-of-the-art results on the popular SUNCG benchmark.
Haiyan Wang 0019, Xuejian Rong, Yingli Tian
ICPR3
2020 Unambiguous Scene Text Segmentation With Referring Expression Comprehension
abstract
Text instance provides valuable information for the understanding and interpretation of natural scenes. The rich, precise high-level semantics embodied in the text could be beneficial for understanding the world around us, and empower a wide range of real-world applications. While most recent visual phrase grounding approaches focus on general objects, this paper explores extracting designated texts and predicting unambiguous scene text segmentation mask, i.e. scene text segmentation from natural language descriptions (referring expressions) like orange text on a little boy in black swinging a bat. The solution of this novel problem enables accurate segmentation of scene text instances from the complex background. In our proposed framework, a unified deep network jointly models visual and linguistic information by encoding both region-level and pixel-level visual features of natural scene images into spatial feature maps, and then decode them into saliency response map of text instances. To conduct quantitative evaluations, we establish a new scene text referring expression segmentation dataset: COCO-CharRef. Experimental results demonstrate the effectiveness of the proposed framework on the text instance segmentation task. By combining image-based visual features with language-based textual explanations, our framework outperforms baselines that are derived from state-of-the-art text localization and natural language object retrieval methods on COCO-CharRef dataset.
Xuejian Rong, Chucai Yi, Yingli Tian
IEEE Trans. Image Process.1
2019 Towards Weakly Supervised Semantic Segmentation in 3D Graph-Structured Point Clouds of Wild Scenes
Haiyan Wang 0019, Xuejian Rong, Shuihua Wang, Yingli Tian
BMVC2
2019 Towards Accurate Instance-Level Text Spotting with Guided Attention
abstract
We tackle the text detection problem from the instance-aware segmentation perspective, in which text bounding boxes are directly extracted from segmentation results without location regression. Specifically, a text-specific attention model and a global enhancement block are introduced to enrich the semantics of text detection features. The attention model is trained with a weakly segmentation supervision signal and enforces the detector to focus on the text regions, while also suppressing the influence of neighboring background clutters. In conjunction with the attention model, a global enhancement block (GEB) is adapted to reason the relationship among different channels with channel-wise weights calibration. Our method achieves comparable performance with the recent state-of-the-arts on ICDAR2013, ICDAR2015, and ICDAR2017-MLT benchmark datasets.
Haiyan Wang 0019, Xuejian Rong, Yingli Tian
ICME2
2019 Incremental Scene Synthesis
abstract
We present a method to incrementally generate complete 2D or 3D scenes with the following properties: (a) it is globally consistent at each step according to a learned scene prior, (b) real observations of a scene can be incorporated while observing global consistency, (c) unobserved regions can be hallucinated locally in consistence with previous observations, hallucinations and global priors, and (d) hallucinations are statistical in nature, i.e., different scenes can be generated from the same observations. To achieve this, we model the virtual scene, where an active agent at each step can either perceive an observed part of the scene or generate a local hallucination. The latter can be interpreted as the agent's expectation at this step through the scene and can be applied to autonomous navigation. In the limit of observing real data at each point, our method converges to solving the SLAM problem. It can otherwise sample entirely imagined scenes from prior distributions. Besides autonomous agents, applications include problems where large data is required for building robust real-world applications, but few samples are available. We demonstrate efficacy on various 2D as well as 3D data.
Benjamin Planche, Xuejian Rong, Ziyan Wu 0001, Srikrishna Karanam, Harald Kosch, Yingli Tian, Jan Ernst, Andreas Hutter
NeurIPS2
2019 Vision-Based Mobile Indoor Assistive Navigation Aid for Blind People
abstract
This paper presents a new holistic vision-based mobile assistive navigation system to help blind and visually impaired people with indoor independent travel. The system detects dynamic obstacles and adjusts path planning in real-time to improve navigation safety. First, we develop an indoor map editor to parse geometric information from architectural models and generate a semantic map consisting of a global 2D traversable grid map layer and context-aware layers. By leveraging the visual positioning service (VPS) within the Google Tango device, we design a map alignment algorithm to bridge the visual area description file (ADF) and semantic map to achieve semantic localization. Using the on-board RGB-D camera, we develop an efficient obstacle detection and avoidance approach based on a time-stamped map Kalman filter (TSM-KF) algorithm. A multi-modal human-machine interface (HMI) is designed with speech-audio interaction and robust haptic interaction through an electronic SmartCane. Finally, field experiments by blindfolded and blind subjects demonstrate that the proposed system provides an effective tool to help blind individuals with indoor navigation and wayfinding.
Bing Li 0008, Juan Pablo Muñoz, Xuejian Rong, Qingtian Chen, Jizhong Xiao, Yingli Tian, Aries Arditi, Mohammed Yousuf
IEEE Trans. Mob. Comput.3
2017 Unambiguous Text Localization and Retrieval for Cluttered Scenes
abstract
Text instance as one category of self-described objects provides valuable information for understanding and describing cluttered scenes. In this paper, we explore the task of unambiguous text localization and retrieval, to accurately localize a specific targeted text instance in a cluttered image given a natural language description that refers to it. To address this issue, first a novel recurrent Dense Text Localization Network (DTLN) is proposed to sequentially decode the intermediate convolutional representations of a cluttered scene image into a set of distinct text instance detections. Our approach avoids repeated detections at multiple scales of the same text instance by recurrently memorizing previous detections, and effectively tackles crowded text instances in close proximity. Second, we propose a Context Reasoning Text Retrieval (CRTR) model, which jointly encodes text instances and their context information through a recurrent network, and ranks localized text bounding boxes by a scoring function of context compatibility. Quantitative evaluations on standard scene text localization benchmarks and a newly collected scene text retrieval dataset demonstrate the effectiveness and advantages of our models for both scene text localization and retrieval.
Xuejian Rong, Chucai Yi, Yingli Tian
CVPR1
2017 Evaluation of Low-Level Features for Real-World Surveillance Event Detection
abstract
Event detection targets at recognizing and localizing specified spatio-temporal patterns in videos. Most research of human activity recognition in the past decades experimented on relatively clean scenes with limited actors performing explicit actions. Recently, more efforts have been paid to the real-world surveillance videos in which the human activity recognition is more challenging due to large variations caused by factors, such as scaling, resolution, viewpoint, cluttered background, and crowdedness. In this paper, we systematically evaluate seven different types of low-level spatio-temporal features in the context of surveillance event detection (SED) using a uniform experimental setup. Fisher vector is employed to aggregate low-level features as the representation of each video clip. A set of random forests is then learned as the classification models. To bridge the research efforts and real-world applications, we utilize the NIST TRECVID SED as our testbed in which seven events are predefined involving different levels of human activity analysis. Strengths and limitations for each low-level feature type are analyzed and discussed.
Yang Xian, Xuejian Rong, Xiaodong Yang 0001, Yingli Tian
IEEE Trans. Circuits Syst. Video Technol.2
2016 Demo: Assisting Visually Impaired People Navigate Indoors
Juan Pablo Muñoz, Bing Li 0008, Xuejian Rong, Jizhong Xiao, Yingli Tian, Aries Arditi
IJCAI3
2016 Region Trajectories for Video Semantic Concept Detection
abstract
Recently, with the advent of the convolutional neural network (CNN), many CNN-based object detection algorithms have been proposed and achieved encouraging results. In this paper, we introduce an algorithm based on region trajectories to establish the connections between object localizations in individual frames and video sequences. To detect object regions in the individual frames of a video, we enhance the region-based convolutional neural network (R-CNN), by incorporating EdgeBox with the Selective Search to generate candidate region proposals and combining the GoogLeNet with the AlexNet to improve the discriminability of the feature representations. The DeepMatching algorithm is employed in our proposed region trajectory method to track the points in the detected object regions. The experiments are conducted on the validation split of the TRECVID 2015 Localization dataset. As demonstrated by the experimental results, our proposed approach improves the object detection accuracy in both temporal and spatial measurements.
Yuancheng Ye, Xuejian Rong, Xiaodong Yang 0001, Yingli Tian
ICMR2
2014 Scene text recognition in multiple frames based on text tracking
abstract
Text signage as visual indicators in natural scene plays an important role in navigation and notification in our daily life. Most previous methods of scene text extraction are developed from a single scene image. In this paper, we propose a multi-frame based scene text recognition method by tracking text regions in a video captured by a moving camera. The main contributions of this paper are as follows. First, we present a framework of scene text recognition in multiple frames based on feature representation of scene text character (STC) for character prediction and conditional random field (CRF) model for word configuration. Second, a feature representation of STC is employed from dense sampled SIFT descriptors and Fisher Vector. Third, we collect a dataset for text information extraction from natural scene videos. Our proposed multi-frame scene text recognition is more compatible with image/video-based mobile applications. The experimental results demonstrate that STC prediction and word configuration in multiple frames based on text tracking significantly improves the performance of scene text recognition.
Xuejian Rong, Chucai Yi, Xiaodong Yang 0001, Yingli Tian
ICME1