EDBT 2026 Demo / reviewers in the wild / expert
Junyu Han
dblp:62/10697
· DBLP profile ↗
52ranked-venue papers
3as first author
32since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 44 · 1 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 40 · 2 first-author · 25 since 2021Databases, data management, data science and information retrieval · 4Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DriveGen: Shared Video-Condition Encoding for Autonomous Multi-View Video GenerationabstractCorner cases, such as severe weather and abnormal lighting, present significant challenges in autonomous driving. The main obstacles involve large-scale data collection and costly annotations. Leveraging generative models to expand corner-case data based on existing annotations offers a promising solution. Unlike monocular videos, multi-view videos introduce an additional "view" dimension, increasing the consistency requirements and making precise control of annotations more challenging. Existing methods decouple multi-view videos along the temporal and view-spatial axes, using separate attention mechanisms, which causes motion discrepancies and limits consistency. Additionally, current approaches employ an independent adapter or ControlNet to encode different 3D annotations, leading to high computational costs and suboptimal alignment between annotations and video latents. These issues arise from neglecting the temporal-spatial relationship and insufficient alignment between 3D annotations and video latents. To address these challenges, we propose DriveGen, which uses 4D position embeddings to encode the positional information of multi-view videos. DriveGen also designs Dual-Scale Full Attention to ensure both global and local spatiotemporal consistency. Furthermore, our Shared Video-Condition Encoding (SVCE) Mechanism converts 3D annotations into 2D masks and encodes both video and annotation sequences using a 3D VAE, requiring only 0.37 M learnable parameters to achieve pixel-level alignment and improving generation quality. Numerous experiments have proven that DriveGen has reached the state-of-the-art, capable of generating high-quality controlled autonomous driving videos. Yuhao Kang, Sanyuan Zhao, Xiameng Qin, Junyu Han, Ji Tao |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2025 | World Knowledge-Enhanced Reasoning Using Instruction-Guided Interactor in Autonomous DrivingabstractThe Multi-modal Large Language Models (MLLMs) with extensive world knowledge have revitalized autonomous driving, particularly in reasoning tasks within perceivable regions. However, when faced with perception-limited areas (dynamic or static occlusion regions), MLLMs struggle to effectively integrate perception ability with world knowledge for reasoning. These perception-limited regions can conceal crucial safety information, especially for vulnerable road users. In this paper, we propose a framework, which aims to improve autonomous driving performance under perception-limited conditions by enhancing the integration of perception capabilities and world knowledge. Specifically, we propose a plug-and-play instruction-guided interaction module that bridges modality gaps and significantly reduces the input sequence length, allowing it to adapt effectively to multi-view video inputs. Furthermore, to better integrate world knowledge with driving-related tasks, we have collected and refined a large-scale multi-modal dataset that includes 2 million natural language QA pairs, 1.7 million grounding task data. To evaluate the model’s utilization of world knowledge, we introduce an object-level risk assessment dataset comprising 200K QA pairs, where the questions necessitate multi-step reasoning leveraging world knowledge for resolution. Extensive experiments validate the effectiveness of our proposed method. Mingliang Zhai, Zengyuan Guo, Ningrui Yang, Xiameng Qin, Sanyuan Zhao, Junyu Han, Ji Tao, Yuwei Wu 0001, Yunde Jia |
AAAI | 7 |
| 2024 | Multi-Domain Incremental Learning for Face Presentation Attack DetectionabstractPrevious face Presentation Attack Detection (PAD) methods aim to improve the effectiveness of cross-domain tasks. However, in real-world scenarios, the original training data of the pre-trained model is not available due to data privacy or other reasons. Under these constraints, general methods for fine-tuning single-target domain data may lose previously learned knowledge, leading to a catastrophic forgetting problem. To address these issues, we propose a multi-domain incremental learning (MDIL) method for PAD, which not only learns knowledge well from the new domain but also maintains the performance of previous domains stably. Specifically, we propose an adaptive domain-specific experts (ADE) framework based on the vision transformer to preserve the discriminability of previous domains. Furthermore, an asymmetric classifier is designed to keep the output distribution of different classifiers consistent, thereby improving the generalization ability. Extensive experiments show that our proposed method achieves state-of-the-art performance compared to prior methods of incremental learning. Excitingly, under more stringent setting conditions, our method approximates or even outperforms the DA/DG-based methods. Keyao Wang, Haixiao Yue, Ajian Liu 0001, Haocheng Feng, Junyu Han, Errui Ding, Jingdong Wang 0001 |
AAAI | 7 |
| 2024 | KD-DETR: Knowledge Distillation for Detection Transformer with Consistent Distillation Points SamplingabstractDETR is a novel end-to-end transformer architecture object detector, which significantly outperforms classic detectors when scaling up. In this paper, we focus on the compression of DETR with knowledge distillation. While knowledge distillation has been well-studied in classic detectors, there is a lack of researches on how to make it work effectively on DETR. We first provide experimental and theoretical analysis to point out that the main challenge in DETR distillation is the lack of consistent distillation points. Distillation points refer to the corresponding inputs of the predictions for student to mimic, which have different formulations in CNN detector and DETR, and reliable distillation requires sufficient distillation points which are consistent between teacher and student. Based on this observation, we propose the first general knowledge distillation paradigm for DETR (KD-DETR) with consistent distillation points sampling, for both homogeneous and heterogeneous distillation. Specifically, we decouple detection and distillation tasks by introducing a set of specialized object queries to construct distillation points for DETR. We further propose a general-to-specific distillation points sampling strategy to explore the extensibility of KD-DETR. Extensive experiments validate the effectiveness and generalization of KD-DETR. For both single-scale DAB-DETR and multis-scale Deformable DETR and DINO, KD-DETR boost the performance of student model with improvements of 2.6% - 5.2%. We further extend KD-DETR to heterogeneous distillation, and achieves 2.1 % improvement by distilling the knowledge from DINO to Faster R-CNN with ResNet-50, which is comparable with homogeneous distillation methods. Shengzhao Weng, Haixiao Yue, Haocheng Feng, Junyu Han, Errui Ding |
CVPR | 7 |
| 2024 | Decoupled Pseudo-Labeling for Semi-Supervised Monocular 3D Object DetectionabstractWe delve into pseudo-labeling for semi-supervised monocular 3D object detection (SSM30D) and discover two primary issues: a misalignment between the prediction quality of 3D and 2D attributes and the tendency of depth supervision derived from pseudo-labels to be noisy, leading to significant optimization conflicts with other re-liable forms of supervision. To tackle these issues, we introduce a novel decoupled pseudo-labeling (DPL) approach for SSM30D. Our approach features a Decoupled Pseudo-label Generation (DPG) module, designed to efficiently generate pseudo-labels by separately processing 2D and 3D attributes. This module incorporates a unique homography-based method for identifying dependable pseudo-labels in Bird's Eye View (BEV) space, specifically for 3D attributes. Additionally, we present a Depth Gradient Projection (DGP) module to mitigate optimization conflicts caused by noisy depth supervision of pseudo-labels, effectively decoupling the depth gradient and re-moving conflicting gradients. This dual decoupling strat-egy-at both the pseudo-label generation and gradient lev-els-significantly improves the utilization of pseudo-labels in SSM30D. Our comprehensive experiments on the KITTI benchmark demonstrate the superiority of our method over existing approaches. Jiaming Li 0010, Xiangru Lin, Wei Zhang 0197, Xiao Tan 0001, Junyu Han, Errui Ding, Jingdong Wang 0001, Guanbin Li |
CVPR | 6 |
| 2024 | CSDG-FAS: Closed-Space Domain Generalization for Face Anti-spoofing
Keyao Wang, Haixiao Yue, Yanyan Liang 0001, Mouxiao Huang, Junyu Han, Errui Ding, Jingdong Wang 0001 |
Int. J. Comput. Vis. | 7 |
| 2024 | BASL-AD SLAM: A Robust Deep-Learning Feature-Based Visual SLAM System With Adaptive Motion ModelabstractVisual Simultaneous Localization and Mapping (VSLAM) plays an important role in advanced driver assistance systems and autonomous driving. Feature-based VSLAM generates very promising and visually pleasant results due to its robustness and localization precision. However, traditional feature-based VSLAM systems are prone to be degraded or fail when either the environment or the motion of robots is too challenging. To handle these problems, we proposed BASL-AD SLAM. Firstly, we leveraged the robustness of deep learning and designed a binary deep learning-based descriptor to enhance the accuracy of feature detection and matching for SLAM systems in challenging environments. Meanwhile, the real-time performance of the SLAM system can be also guaranteed. Furthermore, we proposed an adaptive motion model to supply more accurate initial poses, which facilitated subsequent feature tracking and pose optimization in SLAM. The performance was validated on public datasets. Results verified that our BASL-AD SLAM can carry out robust feature matching and tracking in real-time under challenging environments, meanwhile, pose estimation accuracy was significantly improved and the proposed SLAM system showed competing robustness and accuracy compared with ORB-SLAM3. Junyu Han, Jiangming Kan |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2023 | Cyclically Disentangled Feature Translation for Face Anti-spoofingabstractCurrent domain adaptation methods for face anti-spoofing leverage labeled source domain data and unlabeled target domain data to obtain a promising generalizable decision boundary. However, it is usually difficult for these methods to achieve a perfect domain-invariant liveness feature disentanglement, which may degrade the final classification performance by domain differences in illumination, face category, spoof type, etc. In this work, we tackle cross-scenario face anti-spoofing by proposing a novel domain adaptation method called cyclically disentangled feature translation network (CDFTN). Specifically, CDFTN generates pseudo-labeled samples that possess: 1) source domain-invariant liveness features and 2) target domain-specific content features, which are disentangled through domain adversarial training. A robust classifier is trained based on the synthetic pseudo-labeled images under the supervision of source domain labels. We further extend CDFTN for multi-target domain adaptation by leveraging data from more unlabeled target domains. Extensive experiments on several public datasets demonstrate that our proposed approach significantly outperforms the state of the art. Code and models are available at https://github.com/vis-face/CDFTN. Haixiao Yue, Keyao Wang, Haocheng Feng, Junyu Han, Errui Ding, Jingdong Wang 0001 |
AAAI | 5 |
| 2023 | Ambiguity-Resistant Semi-Supervised Learning for Dense Object DetectionabstractWith basic Semi-Supervised Object Detection (SSOD) techniques, one-stage detectors generally obtain limited promotions compared with two-stage clusters. We experimentally find that the root lies in two kinds of ambiguities: (1) Selection ambiguity that selected pseudo labels are less accurate, since classification scores cannot properly represent the localization quality. (2) Assignment ambiguity that samples are matched with improper labels in pseudo-label assignment, as the strategy is misguided by missed objects and inaccurate pseudo boxes. To tackle these problems, we propose a Ambiguity-Resistant Semi-supervised Learning (ARSL) for one-stage detectors. Specifically, to alleviate the selection ambiguity, Joint-Confidence Estimation (JCE) is proposed to jointly quantifies the classification and localization quality of pseudo labels. As for the assignment ambiguity, Task-Separation Assignment (TSA) is introduced to assign labels based on pixel-level predictions rather than unreliable pseudo boxes. It employs a ‘divide-and-conquer’ strategy and separately exploits positives for the classification and localization task, which is more robust to the assignment ambiguity. Comprehensive experiments demonstrate that ARSL effectively mitigates the ambiguities and achieves state-of-the-art SSOD performance on MS COCO and PASCAL VOC. Codes can be found at https://github.com/PaddlePaddle/PaddleDetection. Chang Liu 0082, Weiming Zhang 0006, Xiangru Lin, Wei Zhang 0197, Xiao Tan 0001, Junyu Han, Xiaomao Li, Errui Ding, Jingdong Wang 0001 |
CVPR | 6 |
| 2023 | PSVT: End-to-End Multi-Person 3D Pose and Shape Estimation with Progressive Video TransformersabstractExisting methods of multi-person video 3D human Pose and Shape Estimation (PSE) typically adopt a two-stage strategy, which first detects human instances in each frame and then performs single-person PSE with temporal model. However, the global spatio-temporal context among spatial instances can not be captured. In this paper, we propose a new end-to-end multi-person 3D Pose and Shape estimation framework with progressive Video Transformer, termed PSVT. In PSVT, a spatio-temporal encoder (STE) captures the global feature dependencies among spatial objects. Then, spatio-temporal pose decoder (STPD) and shape decoder (STSD) capture the global dependencies between pose queries and feature tokens, shape queries and feature tokens, respectively. To handle the variances of objects as time proceeds, a novel scheme of progressive decoding is used to update pose and shape queries at each frame. Besides, we propose a novel pose-guided attention (PGA) for shape decoder to better predict shape parameters. The two components strengthen the decoder of PSVT to improve performance. Extensive experiments on the four datasets show that PSVT achieves stage-of-the-art results. Zhongwei Qiu, Qiansheng Yang, Jian Wang 0066, Haocheng Feng, Junyu Han, Errui Ding, Chang Xu 0002, Dongmei Fu, Jingdong Wang 0001 |
CVPR | 5 |
| 2023 | Semi-DETR: Semi-Supervised Object Detection with Detection TransformersabstractWe analyze the DETR-based framework on semi-supervised object detection (SSOD) and observe that (1) the one-to-one assignment strategy generates incorrect matching when the pseudo ground-truth bounding box is inaccurate, leading to training inefficiency; (2) DETR-based detectors lack deterministic correspondence between the input query and its prediction output, which hinders the applicability of the consistency-based regularization widely used in current SSOD methods. We present Semi-DETR, the first transformer-based end-to-end semi-supervised object detector, to tackle these problems. Specifically, we propose a Stage-wise Hybrid Matching strategy that combines the one-to-many assignment and one-to-one assignment strategies to improve the training efficiency of the first stage and thus provide high-quality pseudo labels for the training of the second stage. Besides, we introduce a Cross-view Query Consistency method to learn the semantic feature invariance of object queries from different views while avoiding the need to find deterministic query correspondence. Furthermore, we propose a Cost-based Pseudo Label Mining module to dynamically mine more pseudo boxes based on the matching cost of pseudo ground truth bounding boxes for consistency training. Extensive experiments on all SSOD settings of both COCO and Pascal VOC benchmark datasets show that our Semi-DETR method outperforms all state-of-the-art methods by clear margins. Xiangru Lin, Wei Zhang 0197, Xiao Tan 0001, Junyu Han, Errui Ding, Jingdong Wang 0001, Guanbin Li |
CVPR | 6 |
| 2023 | Group DETR: Fast DETR Training with Group-Wise One-to-Many AssignmentabstractDetection transformer (DETR) relies on one-to-one assignment, assigning one ground-truth object to one prediction, for end-to-end detection without NMS post-processing. It is known that one-to-many assignment, assigning one ground-truth object to multiple predictions, succeeds in detection methods such as Faster R-CNN and FCOS. While the naive one-to-many assignment does not work for DETR, and it remains challenging to apply one-to-many assignment for DETR training. In this paper, we introduce Group DETR, a simple yet efficient DETR training approach that introduces a group-wise way for one-to-many assignment. This approach involves using multiple groups of object queries, conducting one-to-one assignment within each group, and performing decoder self-attention separately. It resembles data augmentation with automatically-learned object query augmentation. It is also equivalent to simultaneously training parameter-sharing networks of the same architecture, introducing more supervision and thus improving DETR training. The inference process is the same as DETR trained normally and only needs one group of queries without any architecture modification. Group DETR is versatile and is applicable to various DETR variants. The experiments show that Group DETR signifi-cantly speeds up the training convergence and improves the performance of various DETR-based models. Code will be available at https://github.com/Atten4Vis/GroupDETR. Qiang Chen 0007, Xiaokang Chen, Jian Wang 0066, Shan Zhang 0002, Haocheng Feng, Junyu Han, Errui Ding, Jingdong Wang 0001 |
ICCV | 7 |
| 2023 | CFCG: Semi-Supervised Semantic Segmentation via Cross-Fusion and Contour Guidance SupervisionabstractCurrent state-of-the-art semi-supervised semantic segmentation (SSSS) methods typically adopt pseudo labeling and consistency regularization between multiple learners with different perturbations. Although the performance is desirable, many issues remain: (1) supervisions from a single learner tend to be noisy which causes unreliable consistency regularization (2) existing pixel-wise confidence-score-based reliability measurement causes potential error accumulation as the training proceeds. In this paper, we propose a novel SSSS framework, called CFCG, which combines cross-fusion and contour guidance supervision to tackle these issues. Concretely, we adopt both image-level and feature-level perturbations to expand feature distribution thus pushing the potential limits of consistency regularization. Then, two particular modules are proposed to enable effective semi-supervised learning under heavy coherent perturbations. Firstly, Cross-Fusion Supervision (CFS) mechanism leverages multiple learners to enhance the quality of pseudo labels. Secondly, we introduce an adaptive contour guidance module (ACGM) to effectively identify unreliable spatial regions in pseudo labels. Finally, our proposed CFCG achieves gains of mIoU +1.40%, +0.89% with a single learner and +1.85%, +1.33% by fusion inference on PASCAL VOC 2012 and on Cityscapes respectively under 1/8 protocols, clearly surpassing previous methods and reaching the state-of-the-art. Shuo Li 0012, Weiming Zhang 0006, Wei Zhang 0197, Xiao Tan 0001, Junyu Han, Errui Ding, Jingdong Wang 0001 |
ICCV | 6 |
| 2023 | Gradient-based Sampling for Class Imbalanced Semi-supervised Object DetectionabstractCurrent semi-supervised object detection (SSOD) algorithms typically assume class balanced datasets (PASCAL VOC etc.) or slightly class imbalanced datasets (MS-COCO, etc). This assumption can be easily violated since real world datasets can be extremely class imbalanced in nature, thus making the performance of semi-supervised object detectors far from satisfactory. Besides, the research for this problem in SSOD is severely under-explored. To bridge this research gap, we comprehensively study the class imbalance problem for SSOD under more challenging scenarios, thus forming the first experimental setting for class imbalanced SSOD (CI-SSOD). Moreover, we propose a simple yet effective gradient-based sampling framework that tackles the class imbalance problem from the perspective of two types of confirmation biases. To tackle confirmation bias towards majority classes, the gradient-based reweighting and gradient-based thresholding modules leverage the gradients from each class to fully balance the influence of the majority and minority classes. To tackle the confirmation bias from incorrect pseudo labels of minority classes, the class-rebalancing sampling module resamples unlabeled data following the guidance of the gradient-based reweighting module. Experiments on three proposed sub-tasks, namely MS-COCO, MS-COCO → Object365 and LVIS, suggest that our method outperforms current class imbalanced object detectors by clear margins, serving as a baseline for future research in CI-SSOD. Code will be available at https://github.com/nightkeepers/CI-SSOD. Jiaming Li 0010, Xiangru Lin, Wei Zhang 0197, Xiao Tan 0001, Junyu Han, Errui Ding, Jingdong Wang 0001, Guanbin Li |
ICCV | 6 |
| 2023 | Group Pose: A Simple Baseline for End-to-End Multi-person Pose EstimationabstractIn this paper, we study the problem of end-to-end multi-person pose estimation. State-of-the-art solutions adopt the DETR-like framework, and mainly develop the complex decoder, e.g., regarding pose estimation as keypoint box detection and combining with human detection in ED-Pose [38], hierarchically predicting with pose decoder and joint (keypoint) decoder in PETR [27].We present a simple yet effective transformer approach, named Group Pose. We simply regard K-keypoint pose estimation as predicting a set of N × K keypoint positions, each from a keypoint query, as well as representing each pose with an instance query for scoring N pose predictions.Motivated by the intuition that the interaction, among across-instance queries of different types, is not directly helpful, we make a simple modification to decoder self-attention. We replace single self-attention over all the N × (K + 1) queries with two subsequent group self-attentions: (i) N within-instance self-attention, with each over K keypoint queries and one instance query, and (ii) (K +1) same-type across-instance self-attention, each over N queries of the same type. The resulting decoder removes the interaction among across-instance type-different queries, easing the optimization and thus improving the performance. Experimental results on MS COCO and Crowd-Pose show that our approach without human box supervision is superior to previous methods with complex decoders, and even is slightly better than ED-Pose that uses human box supervision. Paddle1and PyTorch2codes are available. Huan Liu 0030, Qiang Chen 0007, Zichang Tan, Jiang-Jiang Liu 0001, Jian Wang 0066, Xiangbo Su, Xiaolong Li 0001, Junyu Han, Errui Ding, Yao Zhao 0001, Jingdong Wang 0001 |
ICCV | 9 |
| 2023 | Graph Contrastive Learning for Skeleton-based Action Recognition
Xiaohu Huang, Hao Zhou 0039, Jian Wang 0066, Haocheng Feng, Junyu Han, Errui Ding, Jingdong Wang 0001, Xinggang Wang, Wenyu Liu 0001, Bin Feng 0001 |
ICLR | 5 |
| 2023 | StrucTexTv2: Masked Visual-Textual Prediction for Document Image Pre-training
Yuechen Yu, Yulin Li 0004, Chengquan Zhang, Xiaoqiang Zhang 0006, Zengyuan Guo, Xiameng Qin, Junyu Han, Errui Ding, Jingdong Wang 0001 |
ICLR | 8 |
| 2023 | MSAbox: A spatially stable face detectorabstractWe observe an elusive defect in face detectors which is overlooked in existing works. Specifically, we find that face detectors output unstable confidence scores when faces are slightly shifted in position. The confidence scores of the shifted faces can be lower than the detection threshold and result in false negatives. We define this phenomenon as face detection spatial instability. In essence, detectors face the spatial instability problem because they perform badly for challenging object positions. An object position is challenging when it is not sufficiently modeled by any existing anchors. To deal with this problem, we propose Matching Similarity Auxiliary (MSA) box, which consists of three parts: MSA assign, Image Digitization Compensator (IDC), and Soft Smooth L1 Loss (SSL-Loss). Specifically, MSA assign and IDC are designed to dig more challenging samples, which guide the face detectors to enhance detection performance in these extreme cases. SSL-Loss balances the training phase based on the regression efforts of each target. It improves the localization performance even with more challenging positive samples involved. Experiments demonstrate that our MSAbox achieves spatially stable face detection. Kangkang Wang, Ziliang Chen 0005, Bin He 0006, Bi Li 0005, Haocheng Feng, Jingtuo Liu, Junyu Han, Errui Ding |
ICME | 9 |
| 2023 | HAP: Structure-Aware Masked Image Modeling for Human-Centric PerceptionabstractModel pre-training is essential in human-centric perception. In this paper, we first introduce masked image modeling (MIM) as a pre-training approach for this task. Upon revisiting the MIM training strategy, we reveal that human structure priors offer significant potential. Motivated by this insight, we further incorporate an intuitive human structure prior - human parts - into pre-training. Specifically, we employ this prior to guide the mask sampling process. Image patches, corresponding to human part regions, have high priority to be masked out. This encourages the model to concentrate more on body structure information during pre-training, yielding substantial benefits across a range of human-centric perception tasks. To further capture human characteristics, we propose a structure-invariant alignment loss that enforces different masked views, guided by the human part prior, to be closely aligned for the same image. We term the entire method as HAP. HAP simply uses a plain ViT as the encoder yet establishes new state-of-the-art performance on 11 human-centric benchmarks, and on-par result on one dataset. For example, HAP achieves 78.1% mAP on MSMT17 for person re-identification, 86.54% mA on PA-100K for pedestrian attribute recognition, 78.2% AP on MS COCO for 2D pose estimation, and 56.0 PA-MPJPE on 3DPW for 3D pose and shape estimation. Junkun Yuan, Xinyu Zhang 0015, Hao Zhou 0039, Jian Wang 0066, Zhongwei Qiu, Zhiyin Shao, Shaofeng Zhang, Sifan Long 0001, Kun Kuang 0001, Junyu Han, Errui Ding, Lanfen Lin, Fei Wu 0001, Jingdong Wang 0001 |
NeurIPS | 11 |
| 2022 | MobileFaceSwap: A Lightweight Framework for Video Face SwappingabstractAdvanced face swapping methods have achieved appealing results. However, most of these methods have many parameters and computations, which makes it challenging to apply them in real-time applications or deploy them on edge devices like mobile phones. In this work, we propose a lightweight Identity-aware Dynamic Network (IDN) for subject-agnostic face swapping by dynamically adjusting the model parameters according to the identity information. In particular, we design an efficient Identity Injection Module (IIM) by introducing two dynamic neural network techniques, including the weights prediction and weights modulation. Once the IDN is updated, it can be applied to swap faces given any target image or video. The presented IDN contains only 0.50M parameters and needs 0.33G FLOPs per frame, making it capable for real-time video face swapping on mobile phones. In addition, we introduce a knowledge distillation-based method for stable training, and a loss reweighting module is employed to obtain better synthesized results. Finally, our method achieves comparable results with the teacher models and other state-of-the-art methods. Zhibin Hong, Changxing Ding, Zhen Zhu 0006, Junyu Han, Jingtuo Liu, Errui Ding |
AAAI | 5 |
| 2022 | ViSTA: Vision and Scene Text Aggregation for Cross-Modal RetrievalabstractVisual appearance is considered to be the most important cue to understand images for cross-modal retrieval, while sometimes the scene text appearing in images can provide valuable information to understand the visual semantics. Most of existing cross-modal retrieval approaches ignore the usage of scene text information and directly adding this information may lead to performance degradation in scene text free scenarios. To address this issue, we propose a full transformer architecture to unify these cross-modal retrieval scenarios in a single Vision and Scene Text Aggregation framework (ViSTA). Specifically, ViSTA utilizes transformer blocks to directly encode image patches and fuse scene text embedding to learn an aggregated visual representation for cross-modal retrieval. To tackle the modality missing problem of scene text, we propose a novel fusion token based transformer aggregation approach to exchange the necessary scene text information only through the fusion token and concentrate on the most important features in each modality. To further strengthen the visual modality, we develop dual contrastive learning losses to embed both image-text pairs and fusion-text pairs into a common cross-modal space. Compared to existing methods, ViSTA enables to aggregate relevant scene text semantics with visual appearance, and hence improve results under both scene text free and scene text aware scenarios. Experimental results show that ViSTA outperforms other methods by at least 8.4% at Recall@ 1 for scene text aware retrieval task. Compared with state-of-the-art scene text free retrieval methods, ViSTA can achieve better accuracy on Flicker30K and MSCOCO while running at least three times faster during the inference stage, which validates the effectiveness of the proposed framework. Mengjun Cheng, Yipeng Sun, Longchao Wang, Xiongwei Zhu, Jie Chen 0001, Guoli Song, Junyu Han, Jingtuo Liu, Errui Ding, Jingdong Wang 0001 |
CVPR | 8 |
| 2022 | Expressive Talking Head Generation with Granular Audio-Visual ControlabstractGenerating expressive talking heads is essential for creating virtual humans. However, existing one- or few-shot methods focus on lip-sync and head motion, ignoring the emotional expressions that make talking faces realistic. In this paper, we propose the Granularly Controlled Audio-Visual Talking Heads (GC-AVT), which controls lip movements, head poses, and facial expressions of a talking head in a granular manner. Our insight is to decouple the audio-visual driving sources through prior-based pre-processing designs. Detailedly, we disassemble the driving image into three complementary parts including: 1) a cropped mouth that facilitates lip-sync; 2) a masked head that implicitly learns pose; and 3) the upper face which works corporately and complementarily with a time-shifted mouth to contribute the expression. Interestingly, the encoded features from the three sources are integrally balanced through reconstruction training. Extensive experiments show that our method generates expressive faces with not only synced mouth shapes, controllable poses, but precisely animated emotional expressions as well. Borong Liang, Yan Pan 0019, Zhizhi Guo, Hang Zhou 0009, Zhibin Hong, Xiaoguang Han 0001, Junyu Han, Jingtuo Liu, Errui Ding, Jingdong Wang 0001 |
CVPR | 7 |
| 2022 | Few-Shot Head Swapping in the WildabstractThe head swapping task aims at flawlessly placing a source head onto a target body, which is of great importance to various entertainment scenarios. While face swapping has drawn much attention, the task of head swapping has rarely been explored, particularly under the few-shot setting. It is inherently challenging due to its unique needs in head modeling and background blending. In this paper, we present the Head Swapper (HeSer), which achieves few-shot head swapping in the wild through two delicately de-signed modules. Firstly, a Head2Head Aligner is devised to holistically migrate pose and expression information from the target to the source head by examining multi-scale in-formation. Secondly, to tackle the challenges of skin color variations and head-background mismatches in the swapping procedure, a Head2Scene Blender is introduced to si-multaneously modify facial skin color and fill mismatched gaps on the background around the head. Particularly, seamless blending is achieved with the help of a Semantic-Guided Color Reference Creation procedure and a Blending UNet. Extensive experiments demonstrate that the proposed method produces superior head swapping results on a variety of scenes. Changyong Shu, Hemao Wu, Hang Zhou 0009, Jiaming Liu 0003, Zhibin Hong, Changxing Ding, Junyu Han, Jingtuo Liu, Errui Ding, Jingdong Wang 0001 |
CVPR | 7 |
| 2022 | Few-Shot Font Generation by Learning Fine-Grained Local StylesabstractFew-shot font generation (FFG), which aims to generate a new font with a few examples, is gaining increasing attention due to the significant reduction in labor cost. A typical FFG pipeline considers characters in a standard font library as content glyphs and transfers them to a new target font by extracting style information from the reference glyphs. Most existing solutions explicitly disentangle content and style of reference glyphs globally or component-wisely. However, the style of glyphs mainly lies in the local details, i.e. the styles of radicals, components, and strokes together depict the style of a glyph. Therefore, even a single character can contain different styles distributed over spatial locations. In this paper, we propose a new font generation approach by learning 1) the fine-grained local styles from references, and 2) the spatial correspondence between the content and reference glyphs. Therefore, each spatial location in the content glyph can be assigned with the right fine-grained style. To this end, we adopt cross-attention over the representation of the content glyphs as the queries and the representations of the reference glyphs as the keys and values. Instead of explicitly disentangling global or component-wise modeling, the cross-attention mechanism can attend to the right local styles in the reference glyphs and aggregate the reference styles into a fine-grained style representation for the given content glyphs. The experiments show that the proposed method outperforms the state-of-the-art methods in FFG. In particular, the user studies also demonstrate the style consistency of our approach significantly outperforms previous methods. Licheng Tang, Yiyang Cai, Jiaming Liu 0003, Zhibin Hong, Mingming Gong, Minhu Fan, Junyu Han, Jingtuo Liu, Errui Ding, Jingdong Wang 0001 |
CVPR | 7 |
| 2022 | UFO: Unified Feature Optimization
Teng Xi, Yifan Sun 0003, Deli Yu, Bi Li 0005, Nan Peng, Xinyu Zhang 0015, Zhigang Wang 0002, Jian Wang 0066, Haocheng Feng, Junyu Han, Jingtuo Liu, Errui Ding, Jingdong Wang 0001 |
ECCV (26) | 13 |
| 2022 | StyleSwap: Style-Based Generator Empowers Robust Face Swapping
Hang Zhou 0009, Zhibin Hong, Ziwei Liu 0002, Jiaming Liu 0003, Zhizhi Guo, Junyu Han, Jingtuo Liu, Errui Ding, Jingdong Wang 0001 |
ECCV (14) | 7 |
| 2022 | RTFormer: Efficient Design for Real-Time Semantic Segmentation with TransformerabstractRecently, transformer-based networks have shown impressive results in semantic segmentation. Yet for real-time semantic segmentation, pure CNN-based approaches still dominate in this field, due to the time-consuming computation mechanism of transformer. We propose RTFormer, an efficient dual-resolution transformer for real-time semantic segmenation, which achieves better trade-off between performance and efficiency than CNN-based models. To achieve high inference efficiency on GPU-like devices, our RTFormer leverages GPU-Friendly Attention with linear complexity and discards the multi-head mechanism. Besides, we find that cross-resolution attention is more efficient to gather global context information for high-resolution branch by spreading the high level knowledge learned from low-resolution branch. Extensive experiments on mainstream benchmarks demonstrate the effectiveness of our proposed RTFormer, it achieves state-of-the-art on Cityscapes, CamVid and COCOStuff, and shows promising results on ADE20K. Jian Wang 0066, Chenhui Gou, Qiman Wu, Haocheng Feng, Junyu Han, Errui Ding, Jingdong Wang 0001 |
NeurIPS | 5 |
| 2022 | Singular Value Fine-tuning: Few-shot Segmentation requires Few-parameters Fine-tuningabstractFreezing the pre-trained backbone has become a standard paradigm to avoid overfitting in few-shot segmentation. In this paper, we rethink the paradigm and explore a new regime: {\em fine-tuning a small part of parameters in the backbone}. We present a solution to overcome the overfitting problem, leading to better model generalization on learning novel classes. Our method decomposes backbone parameters into three successive matrices via the Singular Value Decomposition (SVD), then {\em only fine-tunes the singular values} and keeps others frozen. The above design allows the model to adjust feature representations on novel classes while maintaining semantic clues within the pre-trained backbone. We evaluate our {\em Singular Value Fine-tuning (SVF)} approach on various few-shot segmentation methods with different backbones. We achieve state-of-the-art results on both Pascal-5$^i$ and COCO-20$^i$ across 1-shot and 5-shot settings. Hopefully, this simple baseline will encourage researchers to rethink the role of backbone fine-tuning in few-shot settings. Yanpeng Sun, Qiang Chen 0007, Jian Wang 0066, Haocheng Feng, Junyu Han, Errui Ding, Jian Cheng 0001, Zechao Li, Jingdong Wang 0001 |
NeurIPS | 6 |
| 2021 | PGNet: Real-time Arbitrarily-Shaped Text Spotting with Point Gathering NetworkabstractThe reading of arbitrarily-shaped text has received increasing research attention. However, existing text spotters are mostly built on two-stage frameworks or character-based methods, which suffer from either Non-Maximum Suppression (NMS), Region-of-Interest (RoI) operations, or character-level annotations. In this paper, to address the above problems, we propose a novel fully convolutional Point Gathering Network (PGNet) for reading arbitrarily-shaped text in real-time. The PGNet is a single-shot text spotter, where the pixel-level character classification map is learned with proposed PG-CTC loss avoiding the usage of character-level annotations. With a PG-CTC decoder, we gather high-level character classification vectors from two-dimensional space and decode them into text symbols without NMS and RoI operations involved, which guarantees high efficiency. Additionally, reasoning the relations between each character and its neighbors, a graph refinement module (GRM) is proposed to optimize the coarse recognition and improve the end-to-end performance. Experiments prove that the proposed method achieves competitive accuracy, meanwhile significantly improving the running speed. In particular, in Total-Text, it runs at 46.7 FPS, surpassing the previous spotters with a large margin. Chengquan Zhang, Fei Qi 0001, Xiaoqiang Zhang 0006, Pengyuan Lv, Junyu Han, Jingtuo Liu, Errui Ding, Guangming Shi |
AAAI | 7 |
| 2021 | FaceController: Controllable Attribute Editing for Face in the WildabstractFace attribute editing aims to generate faces with one or multiple desired face attributes manipulated while other details are preserved. Unlike prior works such as GAN inversion which has an expensive reverse mapping process, we propose a simple feed-forward network to generate high-fidelity manipulated faces. By simply employing some existing and easy-obtainable prior information, our method can control, transfer, and edit diverse attributes of faces in the wild. The proposed method can consequently be applied to various applications such as face swapping, face relighting, and makeup transfer. In our method, we decouple identity, expression, pose, and illumination by using 3D priors; separate texture and colors by using region-wise style codes. All the information is embedded into adversarial learning by our identity-style normalization module. Disentanglement losses are proposed to enhance the generator to extract information independently from each attribute. Comprehensive quantitative and qualitative evaluations have been conducted. In a single framework, our method achieves the best or competitive scores on a variety of face applications. Xiyu Yu, Zhibin Hong, Zhen Zhu 0006, Junyu Han, Jingtuo Liu, Errui Ding, Xiang Bai |
AAAI | 5 |
| 2021 | Dynamic Class Queue for Large Scale Face Recognition in the WildabstractLearning discriminative representation using large-scale face datasets in the wild is crucial for real-world applications, yet it remains challenging. The difficulties lie in many aspects and this work focus on computing resource constraint and long-tailed class distribution. Recently, classification-based representation learning with deep neural networks and well-designed losses have demonstrated good recognition performance. However, the computing and memory cost linearly scales up to the number of identities (classes) in the training set, and the learning process suffers from unbalanced classes. In this work, we propose a dynamic class queue (DCQ) to tackle these two problems. Specifically, for each iteration during training, a subset of classes for recognition are dynamically selected and their class weights are dynamically generated on-the-fly which are stored in a queue. Since only a subset of classes is selected for each iteration, the computing requirement is reduced. By using a single server without model parallel, we empirically verify in large-scale datasets that 10% of classes are sufficient to achieve similar performance as using all classes. Moreover, the class weights are dynamically generated in a few-shot manner and therefore suitable for tail classes with only a few instances. We show clear improvement over a strong baseline in the largest public dataset Megaface Challenge2 (MF2) which has 672K identities and over 88% of them have less than 10 instances. Code is available at https://github.com/bilylee/DCQ Bi Li 0005, Teng Xi, Haocheng Feng, Junyu Han, Jingtuo Liu, Errui Ding, Wenyu Liu 0001 |
CVPR | 5 |
| 2021 | StrucTexT: Structured Text Understanding with Multi-Modal TransformersabstractStructured text understanding on Visually Rich Documents (VRDs) is a crucial part of Document Intelligence. Due to the complexity of content and layout in VRDs, structured text understanding has been a challenging task. Most existing studies decoupled this problem into two sub-tasks: entity labeling and entity linking, which require an entire understanding of the context of documents at both token and segment levels. However, little work has been concerned with the solutions that efficiently extract the structured data from different levels. This paper proposes a unified framework named StrucTexT, which is flexible and effective for handling both sub-tasks. Specifically, based on the transformer, we introduce a segment-token aligned encoder to deal with the entity labeling and entity linking tasks at different levels of granularity. Moreover, we design a novel pre-training strategy with three self-supervised tasks to learn a richer representation. StrucTexT uses the existing Masked Visual Language Modeling task and the new Sentence Length Prediction and Paired Boxes Direction tasks to incorporate the multi-modal information across text, image, and layout. We evaluate our method for structured text understanding at segment-level and token-level and show it outperforms the state-of-the-art counterparts with significantly superior performance on the FUNSD, SROIE, and EPHOIE datasets. Yulin Li 0004, Yuxi Qian, Yuechen Yu, Xiameng Qin, Chengquan Zhang, Junyu Han, Jingtuo Liu, Errui Ding |
ACM Multimedia | 8 |
| 2020 | HAMBox: Delving Into Mining High-Quality Anchors on Face DetectionabstractCurrent face detectors utilize anchors to frame a multi-task learning problem which combines classification and bounding box regression. Effective anchor design and anchor matching strategy enable face detectors to localize faces under large pose and scale variations. However, we observe that, more than 80% correctly predicted bounding boxes are regressed from the unmatched anchors (the IoUs between anchors and target faces are lower than a threshold) in the inference phase. It indicates that these unmatched anchors perform excellent regression ability, but the existing methods neglect to learn from them. In this paper, we propose an Online High-quality Anchor Mining Strategy (HAMBox), which explicitly helps outer faces compensate with high-quality anchors. Our proposed HAMBox method could be a general strategy for anchor-based single-stage face detection. Experiments on various datasets, including WIDER FACE, FDDB, AFW and PASCAL Face, demonstrate the superiority of the proposed method. Yang Liu 0356, Xu Tang 0007, Junyu Han, Jingtuo Liu, Dinger Rui, Xiang Wu 0001 |
CVPR | 3 |
| 2020 | Towards Accurate Scene Text Recognition With Semantic Reasoning NetworksabstractScene text image contains two levels of contents: visual texture and semantic information. Although the previous scene text recognition methods have made great progress over the past few years, the research on mining semantic information to assist text recognition attracts less attention, only RNN-like structures are explored to implicitly model semantic information. However, we observe that RNN based methods have some obvious shortcomings, such as time-dependent decoding manner and one-way serial transmission of semantic context, which greatly limit the help of semantic information and the computation efficiency. To mitigate these limitations, we propose a novel end-to-end trainable framework named semantic reasoning network (SRN) for accurate scene text recognition, where a global semantic reasoning module (GSRM) is introduced to capture global semantic context through multi-way parallel transmission. The state-of-the-art results on 7 public benchmarks, including regular text, irregular text and non-Latin long text, verify the effectiveness and robustness of the proposed method. In addition, the speed of SRN has significant advantages over the RNN based methods, demonstrating its value in practical use. Deli Yu, Chengquan Zhang, Junyu Han, Jingtuo Liu, Errui Ding |
CVPR | 5 |
| 2020 | Learning Global Structure Consistency for Robust Object TrackingabstractFast appearance variations and the distractions of similar objects are two of the most challenging problems in visual object tracking. Unlike many existing trackers that focus on modeling only the target, in this work, we consider the transient variations of the whole scene. The key insight is that the object correspondence and spatial layout of the whole scene are consistent (i.e., global structure consistency) in consecutive frames which helps to disambiguate the target from distractors. Moreover, modeling transient variations enables to localize the target under fast variations. Specifically, we propose an effective and efficient short-term model that learns to exploit the global structure consistency in a short time and thus can handle fast variations and distractors. Since short-term modeling falls short of handling occlusion and out of the views, we adopt the long-short term paradigm and use a long-term model that corrects the short-term model when it drifts away from the target or the target is not present. These two components are carefully combined to achieve the balance of stability and plasticity during tracking. We empirically verify that the proposed tracker can tackle the two challenging scenarios and validate it on large scale benchmarks. Remarkably, our tracker improves state-of-the-art-performance on VOT2018 from 0.440 to 0.460, GOT-10k from 0.611 to 0.640, and NFS from 0.619 to 0.629. Bi Li 0005, Chengquan Zhang, Zhibin Hong, Xu Tang 0007, Jingtuo Liu, Junyu Han, Errui Ding, Wenyu Liu 0001 |
ACM Multimedia | 6 |
| 2019 | Look More Than Once: An Accurate Detector for Text of Arbitrary ShapesabstractPrevious scene text detection methods have progressed substantially over the past years. However, limited by the receptive field of CNNs and the simple representations like rectangle bounding box or quadrangle adopted to describe text, previous methods may fall short when dealing with more challenging text instances, such as extremely long text and arbitrarily shaped text. To address these two problems, we present a novel text detector namely LOMO, which localizes the text progressively for multiple times (or in other word, LOok More than Once). LOMO consists of a direct regressor (DR), an iterative refinement module (IRM) and a shape expression module (SEM). At first, text proposals in the form of quadrangle are generated by DR branch. Next, IRM progressively perceives the entire long text by iterative refinement based on the extracted feature blocks of preliminary proposals. Finally, a SEM is introduced to reconstruct more precise representation of irregular text by considering the geometry properties of text instance, including text region, text center line and border offsets. The state-of-the-art results on several public benchmarks including ICDAR2017-RCTW, SCUT-CTW1500, Total-Text, ICDAR2015 and ICDAR17-MLT confirm the striking robustness and effectiveness of LOMO. Chengquan Zhang, Borong Liang, Zuming Huang, Mengyi En, Junyu Han, Errui Ding, Xinghao Ding |
CVPR | 5 |
| 2019 | Chinese Street View Text: Large-Scale Chinese Text Reading With Partially Supervised LearningabstractMost existing text reading benchmarks make it difficult to evaluate the performance of more advanced deep learning models in large vocabularies due to the limited amount of training data. To address this issue, we introduce a new large-scale text reading benchmark dataset named Chinese Street View Text (C-SVT) with 430,000 street view images, which is at least 14 times as large as the existing Chinese text reading benchmarks. To recognize Chinese text in the wild while keeping large-scale datasets labeling cost-effective, we propose to annotate one part of the C-SVT dataset (30,000 images) in locations and text labels as full annotations and add 400,000 more images, where only the corresponding text-of-interest in the regions is given as weak annotations. To exploit the rich information from the weakly annotated data, we design a text reading network in a partially supervised learning framework, which enables to localize and recognize text, learn from fully and weakly annotated data simultaneously. To localize the best matched text proposals from weakly labeled images, we propose an online proposal matching module incorporated in the whole model, spotting the keyword regions by sharing parameters for end-to-end training. Compared with fully supervised training algorithms, this model can improve the end-to-end recognition performance remarkably by 4.03% in F-score at the same labeling cost. The proposed model can also achieve state-of-the-art results on the ICDAR 2017-RCTW dataset, which demonstrates the effectiveness of the proposed partially supervised learning framework. Yipeng Sun, Jiaming Liu 0003, Wei Liu 0091, Junyu Han, Errui Ding, Jingtuo Liu |
ICCV | 4 |
| 2019 | ACFNet: Attentional Class Feature Network for Semantic SegmentationabstractRecent works have made great progress in semantic segmentation by exploiting richer context, most of which are designed from a spatial perspective. In contrast to previous works, we present the concept of class center which extracts the global context from a categorical perspective. This class-level context describes the overall representation of each class in an image. We further propose a novel module, named Attentional Class Feature (ACF) module, to calculate and adaptively combine different class centers according to each pixel. Based on the ACF module, we introduce a coarse-to-fine segmentation network, called Attentional Class Feature Network (ACFNet), which can be composed of an ACF module and any off-the-shell segmentation network (base network). In this paper, we use two types of base networks to evaluate the effectiveness of ACFNet. We achieve new state-of-the-art performance of 81.85% mIoU on Cityscapes dataset with only finely annotated data used for training. Yanqin Chen, Zhihang Li, Zhibin Hong, Jingtuo Liu, Feifei Ma, Junyu Han, Errui Ding |
ICCV | 7 |
| 2019 | ICDAR2019 Robust Reading Challenge on Arbitrary-Shaped Text - RRC-ArTabstractThis paper reports the ICDAR2019 Robust Reading Challenge on Arbitrary-Shaped Text - RRC-ArT that consists of three major challenges: i) scene text detection, ii) scene text recognition, and iii) scene text spotting. A total of 78 submissions from 46 unique teams/individuals were received for this competition. The top performing score of each challenge is as follows: i) T1 - 82.65%, ii) T2.1 - 74.3%, iii) T2.2 - 85.32%, iv) T3.1 - 53.86%, and v) T3.2 - 54.91%. Apart from the results, this paper also details the ArT dataset, tasks description, evaluation metrics and participants' methods. The dataset, the evaluation kit as well as the results are publicly available at the challenge website. Chee Kheng Chng, Errui Ding, Jingtuo Liu, Dimosthenis Karatzas, Chee Seng Chan, Yipeng Sun, Chun Chet Ng, Canjie Luo, Zihan Ni, ChuanMing Fang, Shuaitao Zhang, Junyu Han |
ICDAR | 14 |
| 2019 | EATEN: Entity-Aware Attention for Single Shot Visual Text ExtractionabstractExtracting Text of Interest (ToI) from images is a crucial part of many OCR applications, such as entity recognition of cards, invoices, and receipts. Most of the existing works employ complicated engineering pipeline, which contains OCR and structure information extraction, to fulfill this task. This paper proposes an Entity-aware Attention Text Extraction Network called EATEN, which is an end-to-end trainable system to extract the ToIs without any post-processing. In the proposed framework, each entity is parsed by its corresponding entity-aware decoder, respectively. Moreover, we innovatively introduce a state transition mechanism which further improves the robustness of visual ToI extraction. In consideration of the absence of public benchmarks, we construct a dataset of almost 0.6 million images in three real-world scenarios (train ticket, passport and business card), which is publicly available at https://github.com/beacandler/EATEN. To the best of our knowledge, EATEN is the first single shot method to extract entities from images. Extensive experiments on these benchmarks demonstrate the state-of-the-art performance of EATEN. He Guo 0003, Xiameng Qin, Jiaming Liu 0003, Junyu Han, Jingtuo Liu, Errui Ding |
ICDAR | 4 |
| 2019 | ICDAR 2019 Competition on Large-Scale Street View Text with Partial Labeling - RRC-LSVTabstractRobust text reading from street view images provides valuable information for various applications. Performance improvement of existing methods in such a challenging scenario heavily relies on the amount of fully annotated training data, which is costly and in-efficient to obtain. To scale up the amount of training data while keeping the labeling procedure cost-effective, this competition introduces a new challenge on Large-scale Street View Text with Partial Labeling (LSVT), providing 5,0000 and 400,000 images in full and weak annotations, respectively. This competition aims to explore the abilities of state-of-the-art methods to detect and recognize text instances from large-scale street view images, closing gaps between research benchmarks and real applications. During the competition period, a total number of 41 teams participate in the two tasks with 132 valid submissions, i.e., text detection and end-to-end text spotting. This paper includes dataset descriptions, task definitions, evaluation protocols and results summaries of ICDAR 2019-LSVT challenge. Yipeng Sun, Dimosthenis Karatzas, Chee Seng Chan, Zihan Ni, Chee Kheng Chng, Canjie Luo, Chun Chet Ng, Junyu Han, Errui Ding, Jingtuo Liu |
ICDAR | 10 |
| 2019 | An End-to-End Video Text Detector with Online TrackingabstractVideo text detection is considered as one of the most difficult tasks in document analysis due to the following two challenges: 1) the difficulties caused by video scenes, i.e., motion blur, illumination changes, and occlusion; 2) the properties of text including variants of fonts, languages, orientations, and shapes. Most existing methods attempt to enhance the performance of video text detection by cooperating with video text tracking, but treat these two tasks separately. In this work, we propose an end-to-end video text detection model with online tracking to address these two challenges. Specifically, in the detection branch, we adopt ConvLSTM to capture spatial structure information and motion memory. In the tracking branch, we convert the tracking problem to text instance association, and an appearance-geometry descriptor with memory mechanism is proposed to generate robust representation of text instances. By integrating these two branches into one trainable framework, they can promote each other and the computational cost is significantly reduced. Experiments on existing video text benchmarks including ICDAR2013 Video, Minetto and YVT demonstrate that the proposed method significantly outperforms state-of-the-art methods. Our method improves F-score by about 2% on all datasets and it can run realtime with 24.36 fps on TITAN Xp. Hongyuan Yu, Chengquan Zhang, Junyu Han, Errui Ding, Liang Wang 0001 |
ICDAR | 4 |
| 2019 | A Single-Shot Arbitrarily-Shaped Text Detector based on Context Attended Multi-Task LearningabstractDetecting scene text of arbitrary shapes has been a challenging task over the past years. In this paper, we propose a novel segmentation-based text detector, namely SAST, which employs a context attended multi-task learning framework based on a Fully Convolutional Network (FCN) to learn various geometric properties for the reconstruction of polygonal representation of text regions. Taking sequential characteristics of text into consideration, a Context Attention Block is introduced to capture long-range dependencies of pixel information to obtain a more reliable segmentation. In post-processing, a Point-to-Quad assignment method is proposed to cluster pixels into text instances by integrating both high-level object knowledge and low-level pixel information in a single shot. Moreover, the polygonal representation of arbitrarily-shaped text can be extracted with the proposed geometric properties much more effectively. Experiments on several benchmarks, including ICDAR2015, ICDAR2017-MLT, SCUT-CTW1500, and Total-Text, demonstrate that SAST achieves better or comparable performance in terms of accuracy. Furthermore, the proposed algorithm runs at 27.63 FPS on SCUT-CTW1500 with a Hmean of 81.0% on a single NVIDIA Titan Xp graphics card, surpassing most of the existing segmentation-based methods. Chengquan Zhang, Fei Qi 0001, Zuming Huang, Mengyi En, Junyu Han, Jingtuo Liu, Errui Ding, Guangming Shi |
ACM Multimedia | 6 |
| 2019 | Editing Text in the WildabstractIn this paper, we are interested in editing text in natural images, which aims to replace or modify a word in the source image with another one while maintaining its realistic look. This task is challenging, as the styles of both background and text need to be preserved so that the edited image is visually indistinguishable from the source image. Specifically, we propose an end-to-end trainable style retention network (SRNet) that consists of three modules: text conversion module, background inpainting module and fusion module. The text conversion module changes the text content of the source image into the target text while keeping the original text style. The background inpainting module erases the original text, and fills the text region with appropriate texture. The fusion module combines the information from the two former modules, and generates the edited text images. To our knowledge, this work is the first attempt to edit text in natural images at the word level. Both visual effects and quantitative results on synthetic and real-world dataset (ICDAR 2013) fully confirm the importance and necessity of modular decomposition. We also conduct extensive experiments to validate the usefulness of our method in various real-world applications such as text image synthesis, augmented reality (AR) translation, information hiding, etc. Chengquan Zhang, Jiaming Liu 0003, Junyu Han, Jingtuo Liu, Errui Ding, Xiang Bai |
ACM Multimedia | 4 |
| 2018 | Detecting Text in the Wild with Deep Character Embedding Network
Jiaming Li 0010, Chengquan Zhang, Yipeng Sun, Junyu Han, Errui Ding |
ACCV (4) | 4 |
| 2018 | TextNet: Irregular Text Reading from Images with an End-to-End Trainable Network
Yipeng Sun, Chengquan Zhang, Zuming Huang, Jiaming Liu 0003, Junyu Han, Errui Ding |
ACCV (3) | 5 |
| 2017 | WordSup: Exploiting Word Annotations for Character Based Text DetectionabstractImagery texts are usually organized as a hierarchy of several visual elements, i.e. characters, words, text lines and text blocks. Among these elements, character is the most basic one for various languages such as Western, Chinese, Japanese, mathematical expression and etc. It is natural and convenient to construct a common text detection engine based on character detectors. However, training character detectors requires a vast of location annotated characters, which are expensive to obtain. Actually, the existing real text datasets are mostly annotated in word or line level. To remedy this dilemma, we propose a weakly supervised framework that can utilize word annotations, either in tight quadrangles or the more loose bounding boxes, for character detector training. When applied in scene text detection, we are thus able to train a robust character detector by exploiting word annotations in the rich large-scale real scene text datasets, e.g. ICDAR15 [19] and COCO-text [39]. The character detector acts as a key role in the pipeline of our text detection engine. It achieves the state-of-the-art performance on several challenging scene text detection benchmarks. We also demonstrate the flexibility of our pipeline by various scenarios, including deformed text detection and math expression recognition. Chengquan Zhang, Yuxuan Luo 0002, Junyu Han, Errui Ding |
ICCV | 5 |
| 2016 | STAR-Net: A SpaTial Attention Residue Network for Scene Text Recognition
Wei Liu 0091, Chaofeng Chen, Kwan-Yee Kenneth Wong, Zhizhong Su, Junyu Han |
BMVC | 5 |
| 2016 | Context-aware mathematical expression recognition: An end-to-end framework and a benchmarkabstractIn this paper we propose a novel end-to-end framework for mathematical expression (ME) recognition. The method uses a convolutional neural network (CNN) to perform mathematical symbol detection and recognition simultaneously incorporating spatial context, and can handle multi-part and touching symbols effectively. To evaluate the performance, we provide a benchmark that contains MEs both from real-life and synthetic data. Images in our dataset undergo multiple variations such as viewpoint, illumination and background. For training, we use pure synthetic data for saving human labeling effort. The proposed method achieved 87% accuracy of total correct for clear images and 45% for cluttered ones. Yuxuan Luo 0002, Junyu Han, Errui Ding, Cheng-Lin Liu 0001 |
ICPR | 5 |
| 2013 | Structure guided fusion for depth map inpainting
Fei Qi 0001, Junyu Han, Pengjin Wang, Guangming Shi, Fu Li 0002 |
Pattern Recognit. Lett. | 2 |
| 2011 | Enhancing Gradient Sparsity for Parametrized Motion EstimationabstractIn this paper, we propose a novel motion estimation framework based on the sparsity associated with gradients of the parametrized motion field. Beginning with Shen and Wu’s sparse model for optic flow estimation [15], we show the sparsity of the motion field can be enhanced by increasing the degree of freedom of the parametrized motion model. With such an enhancement, we formulate the motion estimation as an ‘0 optimization problem. Along with an ‘1 norm regularization to the instant constancy assumption, this problem is solved by a reweighted ‘1 optimization approach. Experiments on constant, pure translational, and affine motion models certify that the enhanced sparsity provides improved accuracy for motion estimation. Junyu Han, Fei Qi 0001, Guangming Shi |
BMVC | 1 |
| 2011 | Gradient sparsity for piecewise continuous optical flow estimationabstractThis paper introduces a new sparse model for robust and reliable optical flow estimation. We show that the sparsity is directly related to gradient fields of optical flow. According to theory on sparse signal recovery, we rigorously formulate the optical flow estimation as an ℓ0optimization problem. Considering the piecewise continuous nature of motion, the basic optical flow constraint is regularized by an ℓ1norm. Then, with convex relaxation, the solution is obtain via the ordinary or reweighted ℓ1optimization approach. Experimental results show the proposed method performs better than traditional methods and deals well with the discontinuities on motion boundaries. Junyu Han, Fei Qi 0001, Guangming Shi |
ICIP | 1 |