VLDB 2026 Research / reviewers in the wild / expert
Youjiang Xu
dblp:183/0069
· DBLP profile ↗
9ranked-venue papers
4as first author
2since 2021 · last 2021
0000-0002-2438-3136ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 2 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Image recognition and object detection · 28% Trustworthy machine learning · 23% Vision and language · 20% |
Topics — the 17 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language
video captioning |
0.9 | 3 | 2018 | Sequential Video VLAD: Training the Aggregation Locally and Temporally · IEEE Trans. Image Process. 2018 Spotting and Aggregating Salient Regions for Video Captioning · ACM Multimedia 2018 Multirate Multimodal Video Captioning · ACM Multimedia 2017 |
Computer vision › Image recognition and object detection › object detection
bounding box refinement |
0.5 | 1 | 2021 | Training Robust Object Detectors From Noisy Category Labels and Imprecise Bounding Boxes · IEEE Trans. Image Process. 2021 |
Machine learning › Trustworthy machine learning › robustness
learning with noisy labels |
0.5 | 1 | 2021 | Faster Meta Update Strategy for Noise-Robust Deep Learning · CVPR 2021 |
Computer vision › Image recognition and object detection
object detection |
0.5 | 1 | 2021 | Training Robust Object Detectors From Noisy Category Labels and Imprecise Bounding Boxes · IEEE Trans. Image Process. 2021 |
Machine learning › Trustworthy machine learning
robustness |
0.5 | 1 | 2021 | Faster Meta Update Strategy for Noise-Robust Deep Learning · CVPR 2021 |
Machine learning › Trustworthy machine learning › robustness › learning with noisy labels
robustness to label noise |
0.5 | 1 | 2021 | Training Robust Object Detectors From Noisy Category Labels and Imprecise Bounding Boxes · IEEE Trans. Image Process. 2021 |
Computer vision › Image recognition and object detection › object detection
robust object detection |
0.5 | 1 | 2021 | Training Robust Object Detectors From Noisy Category Labels and Imprecise Bounding Boxes · IEEE Trans. Image Process. 2021 |
Computer vision › Image recognition and object detection
scene text detection |
0.4 | 1 | 2019 | Geometry Normalization Networks for Accurate Scene Text Detection · ICCV 2019 |
Computer vision › Video understanding and tracking
action recognition |
0.3 | 1 | 2018 | Sequential Video VLAD: Training the Aggregation Locally and Temporally · IEEE Trans. Image Process. 2018 |
Computer vision › Vision and language › video captioning
dense video captioning |
0.3 | 1 | 2018 | Spotting and Aggregating Salient Regions for Video Captioning · ACM Multimedia 2018 |
Machine learning › Representation and self-supervised learning
multimodal representation learning |
0.3 | 1 | 2018 | Movie Question Answering: Remembering the Textual Cues for Layered Visual Contents · AAAI 2018 |
Computer vision › Video understanding and tracking
video question answering |
0.3 | 1 | 2018 | Movie Question Answering: Remembering the Textual Cues for Layered Visual Contents · AAAI 2018 |
Computer vision › Video understanding and tracking
video representation learning |
0.3 | 1 | 2018 | Sequential Video VLAD: Training the Aggregation Locally and Temporally · IEEE Trans. Image Process. 2018 |
Computer vision › Video understanding and tracking
temporal modeling |
0.3 | 1 | 2017 | Multirate Multimodal Video Captioning · ACM Multimedia 2017 |
Machine learning › Transfer learning and domain adaptation › meta-learning
meta-gradient |
0.1 | 1 | 2021 | Faster Meta Update Strategy for Noise-Robust Deep Learning · CVPR 2021 |
Machine learning › Transfer learning and domain adaptation
meta-learning |
0.1 | 1 | 2021 | Faster Meta Update Strategy for Noise-Robust Deep Learning · CVPR 2021 |
Computer vision › Vision and language
multimodal fusion |
0.1 | 1 | 2017 | Multirate Multimodal Video Captioning · ACM Multimedia 2017 |
Methods — techniques the papers use, named apart from their topics
meta-learning · 1.0noisy label learning · 0.5meta-gradient · 0.5layer-wise approximation · 0.5geometry normalization · 0.4data augmentation · 0.4convolutional neural network · 0.4subtitle-text alignment · 0.3memory network · 0.3attention mechanism · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Faster Meta Update Strategy for Noise-Robust Deep LearningabstractIt has been shown that deep neural networks are prone to overfitting on biased training data. Towards addressing this issue, meta-learning employs a meta model for correcting the training bias. Despite the promising performances, super slow training is currently the bottleneck in the meta learning approaches. In this paper, we introduce a novel Faster Meta Update Strategy (FaMUS) to replace the most expensive step in the meta gradient computation with a faster layer-wise approximation. We empirically find that FaMUS yields not only a reasonably accurate but also a low-variance approximation of the meta gradient. We conduct extensive experiments to verify the proposed method on two tasks. We show our method is able to save two-thirds of the training time while still maintaining the comparable or achieving even better generalization performance. In particular, our method achieves the state-of-the-art performance on both synthetic and realistic noisy labels, and obtains promising performance on long-tailed recognition on standard benchmarks. Code are released at https://github.com/youjiangxu/FaMUS. Youjiang Xu, Linchao Zhu, Lu Jiang 0004, Yi Yang 0001 |
CVPR | 1 |
| 2021 | Training Robust Object Detectors From Noisy Category Labels and Imprecise Bounding BoxesabstractObject detection has gained great improvements with the advances of convolutional neural networks and the availability of large amounts of accurate training data. Though the amount of data is increasing significantly, the quality of data annotations is not guaranteed from the existing crowd-sourcing labeling platforms. In addition to noisy category labels, imprecise bounding box annotations are commonly existed for object detection data. When the quality of training data degenerates, the performance of the typical object detectors is severely impaired. In this paper, we propose a Meta-Refine-Net (MRNet) to train object detectors from noisy category labels and imprecise bounding boxes. First, MRNet learns to adaptively assign lower weights to proposals with incorrect labels so as to suppress large loss values generated by these proposals on the classification branch. Second, MRNet learns to dynamically generate more accurate bounding box annotations to overcome the misleading of imprecisely annotated bounding boxes. Thus, the imprecise bounding boxes could impose positive impacts on the regression branch rather than simply be ignored. Third, we propose to refine the imprecise bounding box annotations by jointly learning from both the category and the localization information. By doing this, the approximation of ground-truth bounding boxes is more accurate while the misleading would be further alleviated. Our MRNet is model-agnostic and is capable of learning from noisy object detection data with only a few clean examples (less than 2%). Extensive experiments on PASCAL VOC 2012 and MS COCO 2017 demonstrate the effectiveness and efficiency of our method. Youjiang Xu, Linchao Zhu, Yi Yang 0001, Fei Wu 0001 |
IEEE Trans. Image Process. | 1 |
| 2019 | Geometry Normalization Networks for Accurate Scene Text DetectionabstractLarge geometry (e.g., orientation) variances are the key challenges in the scene text detection. In this work, we first conduct experiments to investigate the capacity of networks for learning geometry variances on detecting scene texts, and find that networks can handle only limited text geometry variances. Then, we put forward a novel Geometry Normalization Module (GNM) with multiple branches, each of which is composed of one Scale Normalization Unit and one Orientation Normalization Unit, to normalize each text instance to one desired canonical geometry range through at least one branch. The GNM is general and readily plugged into existing convolutional neural network based text detectors to construct end-to-end Geometry Normalization Networks (GNNets). Moreover, we propose a geometry-aware training scheme to effectively train the GNNets by sampling and augmenting text instances from a uniform geometry variance distribution. Finally, experiments on popular benchmarks of ICDAR 2015 and ICDAR 2017 MLT validate that our method outperforms all the state-of-the-art approaches remarkably by obtaining one-forward test F-scores of 88.52 and 74.54 respectively. Jiaqi Duan, Youjiang Xu, Zhanghui Kuang, Xiaoyu Yue, Wayne Zhang 0001 |
ICCV | 2 |
| 2018 | Movie Question Answering: Remembering the Textual Cues for Layered Visual ContentsabstractMovies provide us with a mass of visual content as well as attracting stories. Existing methods have illustrated that understanding movie stories through only visual content is still a hard problem. In this paper, for answering questions about movies, we put forward a Layered Memory Network (LMN) that represents frame-level and clip-level movie content by the Static Word Memory module and the Dynamic Subtitle Memory module, respectively. Particularly, we firstly extract words and sentences from the training movie subtitles. Then the hierarchically formed movie representations, which are learned from LMN, not only encode the correspondence between words and visual content inside frames, but also encode the temporal alignment between sentences and frames inside movie clips. We also extend our LMN model into three variant frameworks to illustrate the good extendable capabilities. We conduct extensive experiments on the MovieQA dataset. With only visual content as inputs, LMN with frame-level representation obtains a large performance improvement. When incorporating subtitles into LMN to form the clip-level representation, we achieve the state-of-the-art performance on the online evaluation task of 'Video+Subtitles'. The good performance successfully demonstrates that the proposed framework of LMN is effective and the hierarchically formed movie representations have good potential for the applications of movie question answering. Bo Wang 0011, Youjiang Xu, Yahong Han, Richang Hong |
AAAI | 2 |
| 2018 | Spotting and Aggregating Salient Regions for Video CaptioningabstractTowards an interpretable video captioning process, we target to locate salient regions of video objects along with the sequentially uttering words. This paper proposes a new framework to automatically spot salient regions in each video frame and simultaneously learn a discriminative spatio-temporal representation for video captioning. First, in a Spot Module, we automatically learn the saliency value of each location to separate salient regions from video content as the foreground and the rest as background by two operations of 'hard separation' and 'soft separation', respectively. Then, in an Aggregate Module, to aggregate the foreground/background descriptors into a discriminative spatio-temporal representation, we devise a trainable video VLAD process to learn the aggregation parameters. Finally, we utilize the attention mechanism to decode the spatio-temporal representations of different regions into video descriptions. Experiments on two benchmark datasets demonstrate our method outperforms most of the state-of-the-art methods in terms of [email protected], METEOR and CIDEr metrics for the task of video captioning. Also examples demonstrate our method can successfully utter words to sequentially salient regions of video objects. Huiyun Wang, Youjiang Xu, Yahong Han |
ACM Multimedia | 2 |
| 2018 | Sequential Video VLAD: Training the Aggregation Locally and TemporallyabstractAs characterizing videos simultaneously from spatial and temporal cues has been shown crucial for the video analysis, the combination of convolutional neural networks and recurrent neural networks, i.e., recurrent convolution networks (RCNs), should be a native framework for learning the spatio-temporal video features. In this paper, we develop a novel sequential vector of locally aggregated descriptor (VLAD) layer, named SeqVLAD, to combine a trainable VLAD encoding process and the RCNs architecture into a whole framework. In particular, sequential convolutional feature maps extracted from successive video frames are fed into the RCNs to learn soft spatio-temporal assignment parameters, so as to aggregate not only detailed spatial information in separate video frames but also fine motion information in successive video frames. Moreover, we improve the gated recurrent unit (GRU) of RCNs by sharing the input-to-hidden parameters and propose an improved GRU-RCN architecture named shared GRU-RCN (SGRU-RCN). Thus, our SGRU-RCN has a fewer parameters and a less possibility of overfitting. In experiments, we evaluate SeqVLAD with the tasks of video captioning and video action recognition. Experimental results on Microsoft Research Video Description Corpus, Montreal Video Annotation Dataset, UCF101, and HMDB51 demonstrate the effectiveness and good performance of our method. Youjiang Xu, Yahong Han, Richang Hong, Qi Tian 0001 |
IEEE Trans. Image Process. | 1 |
| 2017 | Top attention in line with time: A light-weight strategyabstractFor video representation, dense sampling along trajectories or optical flow stacking are both heavy-cost computations. This paper aims to develop a light-weight strategy which could skip the computations of optical flow and trajectories. Particularly, taking frames as inputs to a pre-trained ConvNet, we extract top layers as video feature maps. Instead of trajectory pooling, we directly pooled these feature maps in line with time, which is named Line Pooling. We utilize the proposed Line-pooled Deep-convolutional Descriptors (LDDs) to weight regions with high motion saliency, which turns out to pay attention to actions in line with time. Experiments on UCF101 and HMDB51 demonstrate the efficiency, effectiveness, and promising performance of our method. Youjiang Xu, Shichao Zhao, Yahong Han, Qinghua Hu, Fei Wu 0001 |
ICME | 1 |
| 2017 | Multirate Multimodal Video CaptioningabstractAutomatically describing videos with natural language is a crucial challenge of video understanding. Compared to images, videos have specific spatial-temporal structure and various modality information. In this paper, we propose a Multirate Multimodal Approach for video captioning. Considering that the speed of motion in videos varies constantly, we utilize a Multirate GRU to capture temporal structure of videos. It encodes video frames with different intervals and has a strong ability to deal with motion speed variance. As videos contain different modality cues, we design a particular multimodal fusion method. By incorporating visual, motion, and topic information together, we construct a well-designed video representation. Then the video representation is fed into a RNN-based language model for generating natural language descriptions. We evaluate our approach for video captioning on "Microsoft Research - Video to Text" (MSR-VTT), a large-scale video benchmark for video understanding. And our approach gets great performance on the 2nd MSR Video to Language Challenge. Ziwei Yang 0001, Youjiang Xu, Huiyun Wang, Bo Wang 0011, Yahong Han |
ACM Multimedia | 2 |
| 2016 | Large-Scale E-Commerce Image Retrieval with Top-Weighted Convolutional Neural NetworksabstractSeveral recent researches have shown that image features produced by Convolutional Neural Networks (CNNs) provide the state-of-the-art performance for image classification and retrieval. Moreover, some researchers have found that the features extracted from the deep convolutional layers of CNNs perform better than that from the fully-connected layers. Features extracted from the convolutional layers have a natural interpretation: descriptors of local image regions correspond well to the receptive fields of the particular features. In order to obtain both representative and discriminative descriptors for large-scale e-commerce image retrieval, we come up with a new feature extraction framework. At first, we propose the Top-Weight method to detect the interesting area of e-commerce images automatically. With the estimated weight, we then aggregate local deep features and produce high-quality global representation for e-commerce image retrieval. We have conducted experiments on an e-commerce dataset ALISC [1] released by Alibaba Group. Experimental results show that our method outperforms other deep learning based methods. Shichao Zhao, Youjiang Xu, Yahong Han |
ICMR | 2 |