Ding Liang

dblp:29/8957 · DBLP profile ↗
← Back
60ranked-venue papers
9as first author
35since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 37 · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 36 · 25 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 9 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2
YearPublicationVenuePosition
2026 DetailGen3D: Generative 3D Geometry Enhancement via Data-Dependent Flow
abstract
Modern 3D generation methods can rapidly create shapes from sparse or single views, but their outputs often lack geometric detail due to computational constraints. We present DetailGen3D, a generative approach specifically designed to enhance these generated 3D shapes. Our key insight is to model the coarse-to-fine transformation directly through data-dependent flows in latent space, avoiding the computational overhead of large-scale 3D generative models. We introduce a token matching strategy that ensures accurate spatial correspondence during refinement, enabling local detail synthesis while preserving global structure. By carefully designing our training data to match the characteristics of synthesized coarse shapes, our method can effectively enhance shapes produced by various 3D generation and reconstruction approaches, from single-view to sparse multi-view inputs. Extensive experiments demonstrate that DetailGen3D achieves high-fidelity geometric detail synthesis while maintaining efficiency in training. Our project page is https://detailgen3d.github.io/DetailGen3D/
Ken Deng, Jingxiang Sun, Zixin Zou, Yangguang Li 0001, Yan-Pei Cao 0001, Yebin Liu, Ding Liang
3DV9
2026 A renaissance of explicit motion information mining from transformers for action recognition
Peiqin Zhuang, Lei Bai 0001, Yichao Wu, Ding Liang, Luping Zhou, Yali Wang 0001, Wanli Ouyang
Pattern Recognit.4
2026 Nexus: Native Mesh Generation with Diffusion
abstract
Generating high-quality triangle meshes is essential for film, gaming, and interactive 3D applications. Mainstream methods rely on mesh serialization and autoregressive processes, which stuggles in effective inference and is sensitive to error accumulation. In this paper, we present Nexus , a diffusion method that achieves holistic mesh generation via decoupled vertex and topology generation. First, we view mesh vertices as sparse voxels organized as an octree and adopt a diffusion model to generate the vertices in a coarse-to-fine manner. Second, for topology modeling, we propose Space-time Interval , as an extension of Spacetime Distance to encode arbitrary edge and face topology into continuous per-vertex embeddings. It allows for a global and efficient recovery of complex topology. We then employ a diffusion model to generate the continuous embeddings on the generated vertices. Extensive experiments on the Objaverse and Toys4K datasets and in-the-wild images demonstrate that our method outperforms state-of-the-art autoregressive and two-stage baselines, effectively circumventing the inherent limitations of sequential mesh modeling. A blind user study from 3D practitioners confirms strong perceptual preference for our results.
Ying-Tian Liu, Qi-Yuan Feng, Zixin Zou, Ding Liang, Biao Zhang 0005, Yan-Pei Cao 0001
ACM Trans. Graph.6
2025 MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation
abstract
This paper introduces MIDI, a novel paradigm for compositional 3D scene generation from a single image. Unlike existing methods that rely on reconstruction or retrieval techniques or recent approaches that employ multi-stage object-by-object generation, MIDI extends pre-trained image-to-3D object generation models to multi-instance diffusion models, enabling the simultaneous generation of multiple 3D instances with accurate spatial relationships and high generalizability. At its core, MIDI incorporates a novel multi-instance attention mechanism, that effectively captures inter-object interactions and spatial coherence directly within the generation process, without the need for complex multi-step processes. The method utilizes partial object images and global scene context as inputs, directly modeling object completion during 3D generation. During training, we effectively supervise the interactions between 3D instances using a limited amount of scene-level data, while incorporating single-object data for regularization, thereby maintaining the pre-trained generalization ability. MIDI demonstrates state-of-the-art performance in image-to-scene generation, validated through evaluations on synthetic data, real-world scene data, and stylized scene images generated by text-to-image diffusion models.
Zehuan Huang, Xingqiao An, Yunhan Yang, Yangguang Li 0001, Zixin Zou, Ding Liang, Xihui Liu, Yan-Pei Cao 0001, Lu Sheng
CVPR7
2025 SparseFlex: High-Resolution and Arbitrary-Topology 3D Shape Modeling
abstract
Creating high-fidelity 3D meshes with arbitrary topology, including open surfaces and complex interiors, remains a significant challenge. Existing implicit field methods often require costly and detail-degrading watertight conversion, while other approaches struggle with high resolutions. This paper introduces SparseFlex, a novel sparse-structured isosurface representation that enables differentiable mesh reconstruction at resolutions up to $1024^3$ directly from rendering losses. SparseFlex combines the accuracy of Flexicubes with a sparse voxel structure, focusing computation on surface-adjacent regions and efficiently handling open surfaces. Crucially, we introduce a frustum-aware sectional voxel training strategy that activates only relevant voxels during rendering, dramatically reducing memory consumption and enabling high-resolution training. This also allows, for the first time, the reconstruction of mesh interiors using only rendering supervision. Building upon this, we demonstrate a complete shape modeling pipeline by training a variational autoencoder (VAE) and a rectified flow transformer for high-quality 3D shape generation. Our experiments show state-of-the-art reconstruction accuracy, with a ~82% reduction in Chamfer Distance and a ~88% increase in F-score compared to previous methods, and demonstrate the generation of high-resolution, detailed 3D shapes with arbitrary topology. By enabling high-resolution, differentiable mesh reconstruction and generation with rendering losses, SparseFlex significantly advances the state-of-the-art in 3D shape representation and modeling.
Xianglong He, Zixin Zou, Chia-Hao Chen, Ding Liang, Chun Yuan 0003, Wanli Ouyang, Yan-Pei Cao 0001, Yangguang Li 0001
ICCV5
2025 NeuFrameQ: Neural Frame Fields for Scalable and Generalizable Anisotropic Quadrangulation
Ying-Tian Liu, Xin Yu 0004, Yan-Pei Cao 0001, Ding Liang, Ariel Shamir, Song-Hai Zhang
ICCV7
2025 ShapeGen: Towards High-Quality 3D Shape Synthesis
abstract
Inspired by generative paradigms in image and video, 3D shape generation has made notable progress, enabling the rapid synthesis of high-fidelity 3D assets from a single image. However, current methods still face challenges, including the lack of intricate details, overly smoothed surfaces, and fragmented thin-shell structures. These limitations leave the generated 3D assets still one step short of meeting the standards favored by artists. In this paper, we present ShapeGen, which achieves high-quality image-to-3D shape generation through 3D representation and supervision improvements, resolution scaling up, and the advantages of linear transformers. These advancements allow the generated assets to be seamlessly integrated into 3D pipelines, facilitating their widespread adoption across various applications. Specifically, in contrast to existing methods: 1) We investigate how different representations and VAE supervision strategies affect the generation process, and address issues like aliasing artifacts and fragmented thin-shell structures by using an TSDF-based representation supervised with BCE loss. 2) We scale up the resolution of 3D data, image conditioning inputs, and the number of latent tokens to enhance generation fidelity. 3) We adopt mixed conditioning using raw RGB images and normal maps during training, effectively resolving ambiguities caused by inconsistencies between ControlNet-generated RGB images and the underlying geometry from untextured assets. 4) We replace the original softmax attention with linear attention to improve training and inference efficiency when handling a large number of latent tokens. 5) We introduce an inference-time scaling strategy that enhances generation quality at test time. Through extensive experiments, we validate the impact of these improvements on overall performance. Ultimately, thanks to the synergistic effects of these enhancements, ShapeGen achieves a significant leap in image-to-3D generation, establishing a new state-of-the-art performance.
Yangguang Li 0001, Xianglong He, Zixin Zou, Zexiang Liu, Wanli Ouyang, Ding Liang, Yan-Pei Cao 0001
SIGGRAPH Asia6
2025 LegoACE: Autoregressive Construction Engine for Expressive LEGO® Assemblies
abstract
Automated LEGO® design is challenging due to the extensive variety of LEGO® brick types and the necessity of constructing semantically meaningful models from individually meaningless components. Current automatic LEGO® generation methods face two key challenges: i) They typically rely on explicit modeling of brick connectivity to ensure structural validity. However, this requires extensive manual annotation, which is labor-intensive as the variety of LEGO® primitives increases. This limits training data diversity, restricting the variety of LEGO® bricks that can be effectively utilized. ii) To facilitate learning within neural networks, current methods often employ either volume or text-based descriptions to represent LEGO® models. However, volumetric representations are computationally expensive and hamper large-scale generative training, while text-based approaches rely on large language models and dedicated text-to-brick mapping rules, introducing a semantic gap between language tokens and 3D brick structures.
Hao Xu 0049, Yuqing Zhang 0005, Xinyang Zheng, Xiangjun Tang, Yunhan Yang, Ding Liang, Yingtian Liu, Yan-Pei Cao 0001, Xiaogang Jin 0001
SIGGRAPH Asia8
2025 OmniPart: Part-Aware 3D Generation with Semantic Decoupling and Structural Cohesion
abstract
The creation of 3D assets with explicit, editable part structures is crucial for advancing interactive applications, yet most generative methods produce only monolithic shapes, limiting their utility. We introduce OmniPart, a novel framework for part-aware 3D object generation designed to achieve high semantic decoupling among components while maintaining robust structural cohesion. OmniPart uniquely decouples this complex task into two synergistic stages: (1) an autoregressive structure planning module generates a controllable, variable-length sequence of 3D part bounding boxes, critically guided by flexible 2D part masks that allow for intuitive control over part decomposition without requiring direct correspondences or semantic labels; and (2) a spatially-conditioned rectified flow model, efficiently adapted from a pre-trained holistic 3D generator, synthesizes all 3D parts simultaneously and consistently within the planned layout. Our approach supports user-defined part granularity, precise localization, and enables diverse downstream applications. Extensive experiments demonstrate that OmniPart achieves state-of-the-art performance, paving the way for more interpretable, editable, and versatile 3D content.
Yunhan Yang, Yufan Zhou 0004, Zixin Zou, Ying-Tian Liu, Hao Xu 0049, Ding Liang, Yan-Pei Cao 0001, Xihui Liu
SIGGRAPH Asia8
2025 SeqTex: Generate Mesh Textures in Video Sequence
abstract
Training native 3D texture generative models remains a fundamental yet challenging problem, largely due to the limited availability of large-scale, high-quality 3D texture datasets. This scarcity hinders generalization to real-world scenarios. To address this, most existing methods finetune foundation image generative models to exploit their learned visual priors. However, these approaches typically generate only multi-view images and rely on post-processing to produce UV texture maps—an essential representation in modern graphics pipelines. Such two-stage pipelines often suffer from error accumulation and spatial inconsistencies across the 3D surface. In this paper, we introduce SeqTex, a novel end-to-end framework that leverages the visual knowledge encoded in pretrained video foundation models to directly generate complete UV texture maps. Unlike previous methods that model the distribution of UV textures in isolation, SeqTex reformulates the task as a sequence generation problem, enabling the model to learn the joint distribution of multi-view renderings and UV textures. This design effectively transfers the consistent image-space priors from video foundation models into the UV domain. To further enhance performance, we propose several architectural innovations: a decoupled multi-view and UV branch design, geometry-informed attention to guide cross-domain feature alignment, and adaptive token resolution to preserve fine texture details while maintaining computational efficiency. Together, these components allow SeqTex to fully utilize pretrained video priors and synthesize high-fidelity UV texture maps without the need for post-processing. Extensive experiments show that SeqTex achieves state-of-the-art performance on both image-conditioned and text-conditioned 3D texture generation tasks, with superior 3D consistency, texture-geometry alignment, and real-world generalization. Our project page is https://yuanze1024.github.io/SeqTex/.
Ze Yuan, Xin Yu 0004, Yang-Tian Sun, Yan-Pei Cao 0001, Ding Liang, Xiaojuan Qi 0001
SIGGRAPH Asia6
2024 EpiDiff: Enhancing Multi-View Synthesis via Localized Epipolar-Constrained Diffusion
abstract
Generating multiview images from a single view facilitates the rapid generation of a 3D mesh conditioned on a single image. Recent methods [31] that introduce 3D global representation into diffusion models have shown the potential to generate consistent multiviews, but they have reduced generation speed and face challenges in maintaining generalizability and quality. To address this issue, we propose EpiDiff, a localized interactive multiview diffusion model. At the core of the proposed approach is to insert a lightweight epipolar attention block into the frozen diffusion model, leveraging epipolar constraints to enable cross-view interaction among feature maps of neighboring views. The newly initialized 3D modeling module preserves the original feature distribution of the diffusion model, exhibiting compatibility with a variety of base diffusion models. Experiments show that EpiDiff generates 16 multiview images in just 12 seconds, and it surpasses previous methods in quality evaluation metrics, including PSNR, SSIM and LPIPS. Additionally, EpiDiff can generate a more diverse distribution of views, improving the reconstruction quality from generated multiviews. Please see the project page at huanngzh.github.io/EpiDiff/.
Zehuan Huang, Junting Dong, Yaohui Wang 0001, Yangguang Li 0001, Yan-Pei Cao 0001, Ding Liang, Yu Qiao 0001, Bo Dai 0002, Lu Sheng
CVPR8
2024 Triplane Meets Gaussian Splatting: Fast and Generalizable Single-View 3D Reconstruction with Transformers
abstract
Recent advancements in 3D reconstruction from single images have been driven by the evolution of generative models. Prominent among these are methods based on Score Distillation Sampling (SDS) and the adaptation ofdiffusion models in the 3D domain. Despite their progress, these techniques often face limitations due to slow optimization or rendering processes, leading to extensive training and optimization times. In this paper, we introduce a novel approach for single-view reconstruction that efficiently generates a 3D model from a single image via feed-forward inference. Our method utilizes two transformer-based networks, namely a point decoder and a triplane decoder, to reconstruct 3D objects using a hybrid Triplane-Gaussian intermediate representation. This hybrid representation strikes a balance, achieving a faster rendering speed compared to implicit representations while simultaneously delivering superior rendering quality than explicit representations. The point decoder is designed for generating point clouds from single images, offering an explicit representation which is then utilized by the triplane decoder to query Gaussian features for each point. This design choice addresses the challenges associated with directly regressing explicit 3D Gaussian attributes characterized by their non-structural nature. Subsequently, the 3D Gaussians are decoded by an MLP to enable rapid rendering through splatting. Both decoders are built upon a scalable, transformer-based architecture and have been efficiently trained on large-scale 3D datasets. The evaluations conducted on both synthetic datasets and real-world images demonstrate that our method not only achieves higher quality but also ensures a faster runtime in comparison to previous state-of-the-art techniques. Please see our project page at https://zouzx.github.io/TriplaneGaussian/
Zixin Zou, Yangguang Li 0001, Ding Liang, Yan-Pei Cao 0001, Song-Hai Zhang
CVPR5
2024 UniDream: Unifying Diffusion Priors for Relightable Text-to-3D Generation
Zexiang Liu, Yangguang Li 0001, Youtian Lin, Xin Yu 0004, Sida Peng, Yan-Pei Cao 0001, Xiaojuan Qi 0001, Xiaoshui Huang, Ding Liang, Wanli Ouyang
ECCV (5)9
2024 Text-to-3D with Classifier Score Distillation
abstract
Text-to-3D generation has made remarkable progress recently, particularly with methods based on Score Distillation Sampling (SDS) that leverages pre-trained 2D diffusion models. While the usage of classifier-free guidance is well acknowledged to be crucial for successful optimization, it is considered an auxiliary trick rather than the most essential component. In this paper, we re-evaluate the role of classifier-free guidance in score distillation and discover a surprising finding: the guidance alone is enough for effective text-to-3D generation tasks. We name this method Classifier Score Distillation (CSD), which can be interpreted as using an implicit classification model for generation. This new perspective reveals new insights for understanding existing techniques. We validate the effectiveness of CSD across a variety of text-to-3D tasks including shape generation, texture synthesis, and shape editing, achieving results superior to those of state-of-the-art methods. Our project page is https://xinyu-andy.github.io/Classifier-Score-Distillation
Xin Yu 0004, Yangguang Li 0001, Ding Liang, Song-Hai Zhang, Xiaojuan Qi 0001
ICLR4
2024 TEXGen: a Generative Diffusion Model for Mesh Textures
abstract
While high-quality texture maps are essential for realistic 3D asset rendering, few studies have explored learning directly in the texture space, especially on large-scale datasets. In this work, we depart from the conventional approach of relying on pre-trained 2D diffusion models for testtime optimization of 3D textures. Instead, we focus on the fundamental problem of learning in the UV texture space itself. For the first time, we train a large diffusion model capable of directly generating high-resolution texture maps in a feed-forward manner. To facilitate efficient learning in high-resolution UV spaces, we propose a scalable network architecture that interleaves convolutions on UV maps with attention layers on point clouds. Leveraging this architectural design, we train a 700 million parameter diffusion model that can generate UV texture maps guided by text prompts and single-view images. Once trained, our model naturally supports various extended applications, including text-guided texture inpainting, sparse-view texture completion, and text-driven texture synthesis. The code is available at https://github.com/CVMI-Lab/TEXGen.
Xin Yu 0004, Ze Yuan, Ying-Tian Liu, Yangguang Li 0001, Yan-Pei Cao 0001, Ding Liang, Xiaojuan Qi 0001
ACM Trans. Graph.8
2023 Improving Robust Fariness via Balance Adversarial Training
abstract
Adversarial training (AT) methods are effective against adversarial attacks, yet they introduce severe disparity of accuracy and robustness between different classes, known as the robust fairness problem. Previously proposed Fair Robust Learning (FRL) adaptively reweights different classes to improve fairness. However, the performance of the better-performed classes decreases, leading to a strong performance drop. In this paper, we observed two unfair phenomena during adversarial training: different difficulties in generating adversarial examples from each class (source-class fairness) and disparate target class tendencies when generating adversarial examples (target-class fairness). From the observations, we propose Balance Adversarial Training (BAT) to address the robust fairness problem. Regarding source-class fairness, we adjust the attack strength and difficulties of each class to generate samples near the decision boundary for easier and fairer model learning; considering target-class fairness, by introducing a uniform distribution constraint, we encourage the adversarial example generation process for each class with a fair tendency. Extensive experiments conducted on multiple datasets (CIFAR-10, CIFAR-100, and ImageNette) demonstrate that our BAT can significantly outperform other baselines in mitigating the robust fairness problem (+5-10\% on the worst class accuracy)(Our codes can be found at https://github.com/silvercherry/Improving-Robust-Fairness-via-Balance-Adversarial-Training).
Chunyu Sun, Chenye Xu, Chengyuan Yao, Siyuan Liang 0004, Yichao Wu, Ding Liang, Xianglong Liu 0001, Aishan Liu
AAAI6
2023 Reconstruct Before Summarize: An Efficient Two-Step Framework for Condensing and Summarizing Meeting Transcripts
abstract
Meetings typically involve multiple participants and lengthy conversations, resulting in redundant and trivial content.To overcome these challenges, we propose a two-step framework, Reconstruct before Summarize (RbS), for effective and efficient meeting summarization.RbS first leverages a self-supervised paradigm to annotate essential contents by reconstructing the meeting transcripts.Secondly, we propose a relative positional bucketing (RPB) algorithm to equip (conventional) summarization models to generate the summary.Despite the additional reconstruction process, our proposed RPB significantly compressed the input, leading to faster processing and reduced memory consumption compared to traditional summarization methods.We validate the effectiveness and efficiency of our method through extensive evaluations and analysis.On two meeting summarization datasets, AMI and ICSI, our approach outperforms previous state-of-the-art approaches without relying on large-scale pretraining or expert-grade annotating tools.
Haochen Tan, Han Wu 0004, Wei Shao 0009, Xinyun Zhang 0001, Mingjie Zhan, Zhaohui Hou, Ding Liang, Linqi Song
EMNLP7
2023 ICD-Face: Intra-class Compactness Distillation for Face Recognition
abstract
Knowledge distillation is an effective model compression method to improve the performance of a lightweight student model by transferring the knowledge of a well-performed teacher model, which has been widely adopted in many computer vision tasks, including face recognition (FR). The current FR distillation methods usually utilize the Feature Consistency Distillation (FCD) (e.g., L2distance) on the learned embeddings extracted by the teacher and student models. However, after using FCD, we observe that the intra-class similarities of the student model are lower than the intra-class similarities of the teacher model a lot. Therefore, we propose an effective FR distillation method called ICD-Face by introducing intra-class compactness distillation into the existing distillation framework. Specifically, in ICD-Face, we first propose to calculate the similarity distributions of the teacher and student models, where the feature banks are introduced to construct sufficient and high-quality positive pairs. Then, we estimate the probability distributions of the teacher and student models and introduce the Similarity Distribution Consistency (SDC) loss to improve the intra-class compactness of the student model. Extensive experimental results on multiple benchmark datasets demonstrate the effectiveness of our proposed ICD-Face for face recognition.
Haoyu Qin, Yichao Wu, Ding Liang
ICCV7
2023 Learning Locality and Isotropy in Dialogue Modeling
Han Wu 0004, Haochen Tan, Mingjie Zhan, Gangming Zhao, Shaoqing Lu, Ding Liang, Linqi Song
ICLR6
2023 Visualizing Severe Weather Events Using JPSS ATMS and VIIRS SDR Data within the ICVS Framework
abstract
Over ten-years, the Integrated Calibration and Validation System (ICVS) Long-Term Monitoring (LTM) System has provided near-real time (NRT) monitoring for Joint Polar Satellite System (JPSS) spacecraft and instruments including their on-orbit status and performance and science data product quality [1] - [4]. The ICVS also harnesses JPSS Sensor Data Record (SDR) data to rapidly (with little latency) visualize radiometric features of severe weather events such as hurricanes and volcanos [5] [6]. This study presents two case studies, one depicting the 3-dimensional (3D) atmospheric warm core structure inside Hurricane Ian from the 2022 North Atlantic Hurricane Season and another showing the 3D temperature structures present during the 2021 Heat Dome event by using JPSS ATMS (and VIIRS for hurricane events) SDR and TDR data. More details and images/animations for hurricane events can be found at https://www.star.nesdis.noaa.gov/smcd/sew/index.php.
Banghua Yan, Jingfeng Huang, Warren Dean Porter, Ding Liang, Ninghai Sun, Lihang Zhou, Quanhua (Mark) Liu, Satya Kalluri
IGARSS4
2023 CycleMLP: A MLP-Like Architecture for Dense Visual Predictions
abstract
This article presents a simple yet effective multilayer perceptron (MLP) architecture, namely CycleMLP, which is a versatile neural backbone network capable of solving various tasks of dense visual predictions such as object detection, segmentation, and human pose estimation. Compared to recent advanced MLP architectures such as MLP-Mixer (Tolstikhin et al. 2021), ResMLP (Touvron et al. 2021), and gMLP (Liu et al. 2021), whose architectures are sensitive to image size and are infeasible in dense prediction tasks, CycleMLP has two appealing advantages: 1) CycleMLP can cope with various spatial sizes of images; 2) CycleMLP achieves linear computational complexity with respect to the image size by using local windows. In contrast, previous MLPs have$O(N^{2})$computational complexity due to their full connections in space. 3) The relationship between convolution, multi-head self-attention in Transformer, and CycleMLP are discussed through an intuitive theoretical analysis. We build a family of models that can surpass state-of-the-art MLP and Transformer models e.g., Swin Transformer (Liu et al. 2021), while using fewer parameters and FLOPs. CycleMLP expands the MLP-like models’ applicability, making them versatile backbone networks that achieve competitive results on dense prediction tasks For example, CycleMLP-Tiny outperforms Swin-Tiny by 1.3% mIoU on ADE20 K dataset with fewer FLOPs. Moreover, CycleMLP also shows excellent zero-shot robustness on ImageNet-C dataset.
Shoufa Chen, Enze Xie, Chongjian Ge, Runjian Chen, Ding Liang, Ping Luo 0002
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 Knowledge Distillation for Object Detection via Rank Mimicking and Prediction-Guided Feature Imitation
abstract
Knowledge Distillation (KD) is a widely-used technology to inherit information from cumbersome teacher models to compact student models, consequently realizing model compression and acceleration. Compared with image classification, object detection is a more complex task, and designing specific KD methods for object detection is non-trivial. In this work, we elaborately study the behaviour difference between the teacher and student detection models, and obtain two intriguing observations: First, the teacher and student rank their detected candidate boxes quite differently, which results in their precision discrepancy. Second, there is a considerable gap between the feature response differences and prediction differences between teacher and student, indicating that equally imitating all the feature maps of the teacher is the sub-optimal choice for improving the student's accuracy. Based on the two observations, we propose Rank Mimicking (RM) and Prediction-guided Feature Imitation (PFI) for distilling one-stage detectors, respectively. RM takes the rank of candidate boxes from teachers as a new form of knowledge to distill, which consistently outperforms the traditional soft label distillation. PFI attempts to correlate feature differences with prediction differences, making feature imitation directly help to improve the student's accuracy. On MS COCO and PASCAL VOC benchmarks, extensive experiments are conducted on various detectors with different backbones to validate the effectiveness of our method. Specifically, RetinaNet with ResNet50 achieves 40.4% mAP on MS COCO, which is 3.5% higher than its baseline, and also outperforms previous KD methods.
Xiang Li 0041, Shanshan Zhang 0001, Yichao Wu, Ding Liang
AAAI6
2022 AnchorFace: Boosting TAR@FAR for Practical Face Recognition
abstract
Within the field of face recognition (FR), it is widely accepted that the key objective is to optimize the entire feature space in the training process and acquire robust feature representations. However, most real-world FR systems tend to operate at a pre-defined False Accept Rate (FAR), and the corresponding True Accept Rate (TAR) represents the performance of the FR systems, which indicates that the optimization on the pre-defined FAR is more meaningful and important in the practical evaluation process. In this paper, we call the predefined FAR as Anchor FAR, and we argue that the existing FR loss functions cannot guarantee the optimal TAR under the Anchor FAR, which impedes further improvements of FR systems. To this end, we propose AnchorFace to bridge the aforementioned gap between the training and practical evaluation process for FR. Given the Anchor FAR, AnchorFace can boost the performance of FR systems by directly optimizing the non-differentiable FR evaluation metrics. Specifically, in AnchorFace, we first calculate the similarities of the positive and negative pairs based on both the features of the current batch and the stored features in the maintained online-updating set. Then, we generate the differentiable TAR loss and FAR loss using a soften strategy. Our AnchorFace can be readily integrated into most existing FR loss functions, and extensive experimental results on multiple benchmark datasets demonstrate the effectiveness of AnchorFace.
Haoyu Qin, Yichao Wu, Ding Liang
AAAI4
2022 PseCo: Pseudo Labeling and Consistency Training for Semi-Supervised Object Detection
Xiang Li 0041, Yichao Wu, Ding Liang, Shanshan Zhang 0001
ECCV (9)5
2022 CoupleFace: Relation Matters for Face Recognition Distillation
Haoyu Qin, Yichao Wu, Jinyang Guo 0002, Ding Liang, Ke Xu 0001
ECCV (12)5
2022 OneFace: One Threshold for All
Haoyu Qin, Yichao Wu, Ding Liang, Gangming Zhao, Ke Xu 0001
ECCV (12)5
2022 CycleMLP: A MLP-like Architecture for Dense Prediction
Shoufa Chen, Enze Xie, Chongjian Ge, Runjian Chen, Ding Liang, Ping Luo 0002
ICLR5
2022 DTG-SSOD: Dense Teacher Guidance for Semi-Supervised Object Detection
abstract
The Mean-Teacher (MT) scheme is widely adopted in semi-supervised object detection (SSOD). In MT, sparse pseudo labels, offered by the final predictions of the teacher (e.g., after Non Maximum Suppression (NMS) post-processing), are adopted for the dense supervision for the student via hand-crafted label assignment. However, the "sparse-to-dense'' paradigm complicates the pipeline of SSOD, and simultaneously neglects the powerful direct, dense teacher supervision. In this paper, we attempt to directly leverage the dense guidance of teacher to supervise student training, i.e., the "dense-to-dense'' paradigm. Specifically, we propose the Inverse NMS Clustering (INC) and Rank Matching (RM) to instantiate the dense supervision, without the widely used, conventional sparse pseudo labels. INC leads the student to group candidate boxes into clusters in NMS as the teacher does, which is implemented by learning grouping information revealed in NMS procedure of the teacher. After obtaining the same grouping scheme as the teacher via INC, the student further imitates the rank distribution of the teacher over clustered candidates through Rank Matching. With the proposed INC and RM, we integrate Dense Teacher Guidance into Semi-Supervised Object Detection (termed "DTG-SSOD''), successfully abandoning sparse pseudo labels and enabling more informative learning on unlabeled data. On COCO benchmark, our DTG-SSOD achieves state-of-the-art performance under various labelling ratios. For example, under 10% labelling ratio, DTG-SSOD improves the supervised baseline from 26.9 to 35.9 mAP, outperforming the previous best method Soft Teacher by 1.9 points.
Xiang Li 0041, Yichao Wu, Ding Liang, Shanshan Zhang 0001
NeurIPS5
2022 PVT v2: Improved baselines with Pyramid Vision Transformer
abstract
Transformers have recently lead to encouraging progress in computer vision. In this work, we present new baselines by improving the original Pyramid Vision Transformer (PVT v1) by adding three designs: (i) a linear complexity attention layer, (ii) an overlapping patch embedding, and (iii) a convolutional feed-forward network. With these modifications, PVT v2 reduces the computational complexity of PVT v1 to linearity and provides significant improvements on fundamental vision tasks such as classification, detection, and segmentation. In particular, PVT v2 achieves comparable or better performance than recent work such as the Swin transformer. We hope this work will facilitate state-of-the-art transformer research in computer vision. Code is available at https://github.com/whai362/PVT .
Wenhai Wang, Enze Xie, Xiang Li 0028, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu 0002, Ping Luo 0002, Ling Shao 0001
Comput. Vis. Media6
2022 PAN++: Towards Efficient and Accurate End-to-End Spotting of Arbitrarily-Shaped Text
abstract
Scene text detection and recognition have been well explored in the past few years. Despite the progress, efficient and accurate end-to-end spotting of arbitrarily-shaped text remains challenging. In this work, we propose an end-to-end text spotting framework, termed PAN++, which can efficiently detect and recognize text of arbitrary shapes in natural scenes. PAN++ is based on the kernel representation that reformulates a text line as a text kernel (central region) surrounded by peripheral pixels. By systematically comparing with existing scene text representations, we show that our kernel representation can not only describe arbitrarily-shaped text but also well distinguish adjacent text. Moreover, as a pixel-based representation, the kernel representation can be predicted by a single fully convolutional network, which is very friendly to real-time applications. Taking the advantages of the kernel representation, we design a series of components as follows: 1) a computationally efficient feature enhancement network composed of stacked Feature Pyramid Enhancement Modules (FPEMs); 2) a lightweight detection head cooperating with Pixel Aggregation (PA); and 3) an efficient attention-based recognition head with Masked RoI. Benefiting from the kernel representation and the tailored components, our method achieves high inference speed while maintaining competitive accuracy. Extensive experiments show the superiority of our method. For example, the proposed PAN++ achieves an end-to-end text spotting F-measure of 64.9 at 29.2 FPS on the Total-Text dataset, which significantly outperforms the previous best method. Code will be available at: git.io/PAN.
Wenhai Wang, Enze Xie, Xiang Li 0041, Xuebo Liu 0001, Ding Liang, Zhibo Yang 0003, Tong Lu 0002, Chunhua Shen
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 Characterization and Correction of Intersensor Calibration Convolution Errors Between S-NPP OMPS Nadir Mapper and Metop-B GOME-2
abstract
This article introduces a method to correct intersensor calibration convolution errors that occur in the convolution of spectral response functions (SRFs) between narrow-band and broad-band instruments. By using the intersensor calibration analysis between Ozone Mapping and Profiler Suite (OMPS) Nadir Mapper (NM) and Global Ozone Monitoring Experiment-2 (GOME-2) as an example, the root cause of convolution errors in the intersensor calibration is addressed through direct comparison of OMPS NM SRF and convolved OMPS SRF with GOME-2 SRF. The results reveal that distorted SRF of the narrow-band instrument is the major cause, which appears for GOME-2 at a wide range of channels. The convolution errors in reflectance, which were ignored in previous studies, can be greater than 2% for wavelength shorter than 320 nm and$\sim 0.5$% for wavelengths between 320 and 330 nm. This study thus presents a hybrid convolution error correction method that consists of theoretical approximation of the convolution errors and empirical estimates of residuals due to the deviation of the theoretical approximation from the actual convolution errors. According to the validation through simulation, after applying convolution error correction, the mean convolution errors are less than 0.02%, while the root mean square errors are reduced from more than 0.5% to less than 0.1%. In addition, the correction method is applied to the intersensor calibration radiometric bias assessment between the Meteorological Operational satellite–B (Metop-B) GOME-2 and the Suomi National Polar-orbiting Partnership (S-NPP) OMPS NM. The averaged intersensor calibration reflectance differences are decreased by more than 16% after convolution error correction.
Ding Liang, Banghua Yan, Lawrence E. Flynn
IEEE Trans. Geosci. Remote. Sens.1
2022 Action Recognition With Motion Diversification and Dynamic Selection
abstract
Motion modeling is crucial in modern action recognition methods. As motion dynamics like moving tempos and action amplitude may vary a lot in different video clips, it poses great challenge on adaptively covering proper motion information. To address this issue, we introduce a Motion Diversification and Selection (MoDS) module to generate diversified spatio-temporal motion features and then select the suitable motion representation dynamically for categorizing the input video. To be specific, we first propose a spatio-temporal motion generation (StMG) module to construct a bank of diversified motion features with varying spatial neighborhood and time range. Then, a dynamic motion selection (DMS) module is leveraged to choose the most discriminative motion feature both spatially and temporally from the feature bank. As a result, our proposed method can make full use of the diversified spatio-temporal motion information, while maintaining computational efficiency at the inference stage. Extensive experiments on five widely-used benchmarks, demonstrate the effectiveness of the method and we achieve state-of-the-art performance on Something-Something V1 & V2 that are of large motion variation.
Peiqin Zhuang, Luping Zhou, Lei Bai 0001, Ding Liang, Zhiyong Wang 0001, Yali Wang 0001, Wanli Ouyang
IEEE Trans. Image Process.6
2021 DAM: Discrepancy Alignment Metric for Face Recognition
abstract
The field of face recognition (FR) has witnessed remarkable progress with the surge of deep learning. The effective loss functions play an important role for FR. In this paper, we observe that a majority of loss functions, including the widespread triplet loss and softmax-based cross-entropy loss, embed inter-class (negative) similarity snand intra-class (positive) similarity spinto similarity pairs and optimize to reduce (sn− sp) in the training process. However, in the verification process, existing metrics directly take the absolute similarity between two features as the confidence of belonging to the same identity, which inevitably causes a gap between the training and verification process. To bridge the gap, we propose a new metric called Discrepancy Alignment Metric (DAM) for verification, which introduces the Local Inter-class Discrepancy (LID) for each face image to normalize the absolute similarity score. To estimate the LID of each face image in the verification process, we propose two types of LID Estimation (LIDE) methods, which are reference-based and learning-based estimation methods, respectively. The proposed DAM is plug-and-play and can be easily applied to the most existing methods. Extensive experiments on multiple popular face recognition benchmark datasets demonstrate the effectiveness of our proposed method.
Yudong Wu, Yichao Wu, Chuming Li, Xiaolin Hu 0001, Ding Liang
ICCV6
2021 Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions
abstract
Although convolutional neural networks (CNNs) have achieved great success in computer vision, this work investigates a simpler, convolution-free backbone network use-fid for many dense prediction tasks. Unlike the recently-proposed Vision Transformer (ViT) that was designed for image classification specifically, we introduce the Pyramid Vision Transformer (PVT), which overcomes the difficulties of porting Transformer to various dense prediction tasks. PVT has several merits compared to current state of the arts. (1) Different from ViT that typically yields low-resolution outputs and incurs high computational and memory costs, PVT not only can be trained on dense partitions of an image to achieve high output resolution, which is important for dense prediction, but also uses a progressive shrinking pyramid to reduce the computations of large feature maps. (2) PVT inherits the advantages of both CNN and Transformer, making it a unified backbone for various vision tasks without convolutions, where it can be used as a direct replacement for CNN backbones. (3) We validate PVT through extensive experiments, showing that it boosts the performance of many downstream tasks, including object detection, instance and semantic segmentation. For example, with a comparable number of parameters, PVT+RetinaNet achieves 40.4 AP on the COCO dataset, surpassing ResNet50+RetinNet (36.3 AP) by 4.1 absolute AP (see Figure 2). We hope that PVT could, serre as an alternative and useful backbone for pixel-level predictions and facilitate future research.
Wenhai Wang, Enze Xie, Xiang Li 0028, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu 0002, Ping Luo 0002, Ling Shao 0001
ICCV6
2021 Segmenting Transparent Objects in the Wild with Transformer
abstract
This work presents a new fine-grained transparent object segmentation dataset, termed Trans10K-v2, extending Trans10K-v1, the first large-scale transparent object segmentation dataset. Unlike Trans10K-v1 that only has two limited categories, our new dataset has several appealing benefits. (1) It has 11 fine-grained categories of transparent objects, commonly occurring in the human domestic environment, making it more practical for real-world application. (2) Trans10K-v2 brings more challenges for the current advanced segmentation methods than its former version. Furthermore, a novel Transformer-based segmentation pipeline termed Trans2Seg is proposed. Firstly, the Transformer encoder of Trans2Seg provides the global receptive field in contrast to CNN's local receptive field, which shows excellent advantages over pure CNN architectures. Secondly, by formulating semantic segmentation as a problem of dictionary look-up, we design a set of learnable prototypes as the query of Trans2Seg's Transformer decoder, where each prototype learns the statistics of one category in the whole dataset. We benchmark more than 20 recent semantic segmentation methods, demonstrating that Trans2Seg significantly outperforms all the CNN-based methods, showing the proposed algorithm's potential ability to solve transparent object segmentation.Code is available in https://github.com/xieenze/Trans2Seg.
Enze Xie, Wenjia Wang 0009, Wenhai Wang, Peize Sun, Hang Xu 0004, Ding Liang, Ping Luo 0002
IJCAI6
2020 Online Knowledge Distillation via Collaborative Learning
abstract
This work presents an efficient yet effective online Knowledge Distillation method via Collaborative Learning, termed KDCL, which is able to consistently improve the generalization ability of deep neural networks (DNNs) that have different learning capacities. Unlike existing two-stage knowledge distillation approaches that pre-train a DNN with large capacity as the ''teacher'' and then transfer the teacher's knowledge to another ''student'' DNN unidirectionally (i.e. one-way), KDCL treats all DNNs as ''students'' and collaboratively trains them in a single stage (knowledge is transferred among arbitrary students during collaborative training), enabling parallel computing, fast computations, and appealing generalization ability. Specifically, we carefully design multiple methods to generate soft target as supervisions by effectively ensembling predictions of students and distorting the input images. Extensive experiments show that KDCL consistently improves all the ''students'' on different datasets, including CIFAR-100 and ImageNet. For example, when trained together by using KDCL, ResNet-50 and MobileNetV2 achieve 78.2% and 74.0% top-1 accuracy on ImageNet, outperforming the original results by 1.4% and 2.0% respectively. We also verify that models pre-trained with KDCL transfer well to object detection and semantic segmentation on MS COCO dataset. For instance, the FPN detector is improved by 0.9% mAP.
Qiushan Guo, Xinjiang Wang, Yichao Wu, Ding Liang, Xiaolin Hu 0001, Ping Luo 0002
CVPR5
2020 Rotation Consistent Margin Loss for Efficient Low-Bit Face Recognition
abstract
In this paper, we consider the low-bit quantization problem of face recognition (FR) under the open-set protocol. Different from well explored low-bit quantization on closed-set image classification task, the open-set task is more sensitive to quantization errors (QEs). We redefine the QEs in angular space and disentangle it into class error and individual error. These two parts correspond to inter-class separability and intra-class compactness, respectively. Instead of eliminating the entire QEs, we propose the rotation consistent margin (RCM) loss to minimize the individual error, which is more essential to feature discriminative power. Extensive experiments on popular benchmark datasets such as MegaFace Challenge, Youtube Faces (YTF), Labeled Face in the Wild (LFW) and IJB-C show the superiority of proposed loss in low-bit FR quantization tasks.
Yudong Wu, Yichao Wu, Ruihao Gong, Yuanhao Lv, Ding Liang, Xiaolin Hu 0001, Xianglong Liu 0001
CVPR6
2020 PolarMask: Single Shot Instance Segmentation With Polar Representation
abstract
In this paper, we introduce an anchor-box free and single shot instance segmentation method, which is conceptually simple, fully convolutional and can be used by easily embedding it into most off-the-shelf detection methods. Our method, termed PolarMask, formulates the instance segmentation problem as predicting contour of instance through instance center classification and dense distance regression in a polar coordinate. Moreover, we propose two effective approaches to deal with sampling high-quality center examples and optimization for dense distance regression, respectively, which can significantly improve the performance and simplify the training process. Without any bells and whistles, PolarMask achieves 32.9% in mask mAP with single-model and single-scale training/testing on the challenging COCO dataset. For the first time, we show that the complexity of instance segmentation, in terms of both design and computation complexity, can be the same as bounding box object detection and this much simpler and flexible instance segmentation framework can achieve competitive accuracy. We hope that the proposed PolarMask framework can serve as a fundamental and strong baseline for single shot instance segmentation task.
Enze Xie, Peize Sun, Xiaoge Song, Wenhai Wang, Xuebo Liu 0001, Ding Liang, Chunhua Shen, Ping Luo 0002
CVPR6
2020 AE TextSpotter: Learning Visual and Linguistic Representation for Ambiguous Text Spotting
Wenhai Wang, Xuebo Liu 0001, Xiaozhong Ji, Enze Xie, Ding Liang, Zhibo Yang 0003, Tong Lu 0002, Chunhua Shen, Ping Luo 0002
ECCV (14)5
2020 Scene Text Image Super-Resolution in the Wild
Wenjia Wang 0009, Enze Xie, Xuebo Liu 0001, Wenhai Wang, Ding Liang, Chunhua Shen, Xiang Bai
ECCV (10)5
2020 Face Image Quality Assessment for Model and Human Perception
abstract
Practical face image quality assessment (FIQA) models are trained under the supervision of labeled data, which requires more or less human labor. The human labeled quality scores are consistent with perceptual intuition but laborious. On the other hand, models can be trained with data generated automatically by the recognition models with artificially selected references. However, the recognition scores are sometimes inaccurate, which may give wrong quality scores during FIQA training. In this paper, we propose a labour-saving method for quality scores generation. For the first time, we conduct systematic investigations to show that there exist severe contradictions between different types of target quality, namely distribution gap (DG). To bridge the gap, we propose a novel framework for training FIQA models by combining the merits of data from different sources. In order to make the target score from multiple sources compatible, we design a method called quality distribution alignment (QDA). Meanwhile, to correct the wrong target by recognition models, contradictory samples selection (CSS) is adopted to select samples from the human labeled dataset adaptively. Extensive experiments and analysis on public benchmarks including MegaFace has demonstrated the superiority of our in terms of effectiveness and efficiency.
Yichao Wu, Zhenmao Li, Yudong Wu, Ding Liang
ICPR5
2020 Dynamic Multi-path Neural Network
abstract
Although deeper and larger neural networks have achieved better performance, due to overwhelming burden on computation, they cannot meet the demands of deployment on resource-limited devices. An effective strategy to address this problem is to make use of dynamic inference mechanism, which changes the inference path for different samples at runtime. Existing methods only reduce the depth by skipping an entire specific layer, which may lose important information in this layer. In this paper, we propose a novel method called Dynamic Multipath Neural Network (DMNN), which provides more topology choices in terms of both width and depth on the fly. For better modelling the inference path selection, we further introduce previous state and object category information to guide the training process. Compared to previous dynamic inference techniques, the proposed method is more flexible and easier to incorporate into most modern network architectures. Experimental results on ImageNet and CIFAR-100 demonstrate the superiority of our method on both efficiency and classification accuracy.
Yingcheng Su, Yichao Wu, Ding Liang, Xiaolin Hu 0001
ICPR4
2020 Lifetime Performance Assessment of SNPP OMPS Nadir MAPPER SDR Data Using Simultaneous Nadir Overpass Collocated Observations with Gome-2
abstract
The Nadir Mapper (NM) is one of two nadir sensors of the Ozone Mapping and Profiler Suite (OMPS) that are designed to measure the ultraviolet radiance backscattered by the Earth's atmosphere and surface as well as solar irradiance. This study assesses the lifetime performance of Suomi National Polar-orbiting Partnership (SNPP) satellite NM reflectance data since its launch by using Simultaneous Nadir Overpass (SNO) collocated observations with Global Ozone Monitoring Experiment-2 (GOME-2) spectrometer onboard Meteorological Operational-B (Metop-B) satellite. The study also analyzes the consistency of the NM data quality between SNPP and NOAA-20 satellite using GOME-2 as a transfer.
Ding Liang, Banghua Yan, Ninghai Sun, Lawrence E. Flynn, Chunhui Pan, Trevor Beck
IGARSS1
2020 Companion Guided Soft Margin for Face Recognition
Yingcheng Su, Yichao Wu, Zhenmao Li, Qiushan Guo, Ding Liang, Xiaolin Hu 0001
ECML/PKDD (3)7
2019 R3 Adversarial Network for Cross Model Face Recognition
abstract
In this paper, we raise a new problem, namely cross model face recognition (CMFR), which has considerable economic and social significance. The core of this problem is to make features extracted from different models comparable. However, the diversity, mainly caused by different application scenarios, frequent version updating, and all sorts of service platforms, obstructs interaction among different models and poses a great challenge. To solve this problem, from the perspective of Bayesian modelling, we propose R3Adversarial Network (R3AN) which consists of three paths: Reconstruction, Representation and Regression. We also introduce adversarial learning into the reconstruction path for better performance. Comprehensive experiments on public datasets demonstrate the feasibility of interaction among different models with the proposed framework. When updating the gallery, R3AN conducts the feature transformation nearly 10 times faster than ResNet-101. Meanwhile, the transformed feature distribution is very close to that of target model, and its error rate is incredibly reduced by approximately 75% compared with a naive transformation model. Furthermore, we show that face feature can be deciphered into original face image roughly by the reconstruction path, which may give valuable hints for improving the original face recognition models.
Yichao Wu, Haoyu Qin, Ding Liang, Xuebo Liu 0001
CVPR4
2019 Dynamic Recursive Neural Network
abstract
This paper proposes the dynamic recursive neural network (DRNN), which simplifies the duplicated building blocks in deep neural network. Different from forwarding through different blocks sequentially in previous networks, we demonstrate that the DRNN can achieve better performance with fewer blocks by employing block recursively. We further add a gate structure to each block, which can adaptively decide the loop times of recursive blocks to reduce the computational cost. Since the recursive networks are hard to train, we propose the Loopy Variable Batch Normalization (LVBN) to stabilize the volatile gradient. Further, we improve the LVBN to correct statistical bias caused by the gate structure. Experiments show that the DRNN reduces the parameters and computational cost and while outperforms the original model in term of the accuracy consistently on CIFAR-10 and ImageNet-1k. Lastly we visualize and discuss the relation between image saliency and the number of loop time.
Qiushan Guo, Yichao Wu, Ding Liang, Haoyu Qin
CVPR4
2019 Knowledge Distillation via Route Constrained Optimization
abstract
Distillation-based learning boosts the performance of the miniaturized neural network based on the hypothesis that the representation of a teacher model can be used as structured and relatively weak supervision, and thus would be easily learned by a miniaturized model. However, we find that the representation of a converged heavy model is still a strong constraint for training a small student model, which leads to a higher lower bound of congruence loss. In this work, we consider the knowledge distillation from the perspective of curriculum learning by teacher's routing. Instead of supervising the student model with a converged teacher model, we supervised it with some anchor points selected from the route in parameter space that the teacher model passed by, as we called route constrained optimization (RCO). We experimentally demonstrate this simple operation greatly reduces the lower bound of congruence loss for knowledge distillation, hint and mimicking learning. On close-set classification tasks like CIFAR and ImageNet, RCO improves knowledge distillation by 2.14% and 1.5% respectively. For the sake of evaluating the generalization, we also test RCO on the open-set face recognition task MegaFace. RCO achieves 84.3% accuracy on one-to-million task with only 0.8 M parameters, which push the SOTA by a large margin.
Baoyun Peng, Yichao Wu, Yu Liu 0015, Ding Liang, Xiaolin Hu 0001
ICCV6
2018 FOTS: Fast Oriented Text Spotting With a Unified Network
abstract
Incidental scene text spotting is considered one of the most difficult and valuable challenges in the document analysis community. Most existing methods treat text detection and recognition as separate tasks. In this work, we propose a unified end-to-end trainable Fast Oriented Text Spotting (FOTS) network for simultaneous detection and recognition, sharing computation and visual information among the two complementary tasks. Specifically, RoIRotate is introduced to share convolutional features between detection and recognition. Benefiting from convolution sharing strategy, our FOTS has little computation overhead compared to baseline text detection network, and the joint training method makes our method perform better than these two-stage methods. Experiments on ICDAR 2015, ICDAR 2017 MLT, and ICDAR 2013 datasets demonstrate that the proposed method outperforms state-of-the-art methods significantly, which further allows us to develop the first real-time oriented text spotting system which surpasses all previous state-of-the-art results by more than 5% on ICDAR 2015 text spotting task while keeping 22.6 fps.
Xuebo Liu 0001, Ding Liang, Shi Yan 0007, Dagui Chen, Yu Qiao 0001
CVPR2
2016 Monitoring of Suomi-NPP OMPS calibration parameters and understanding their impacts on earth view radiance
abstract
The key calibration parameters important to instrument health and safety of Ozone Mapping Profiler Suite (OMPS) Nadir Sensors on board Suomi NPP (Suomi National Polar-orbiting Partnership) are being monitored since November 2011 at NOAA/STAR by the Intensive Calibration and Validation System (ICVA). OMPS has two instrument modules: a combined Nadir Mapper (NM) and Nadir Profiler (NP), and a separate Limb Profiler (LP). Both nadir sensors are designed to make measurements of the ultraviolet radiance backscattered by the Earth's atmosphere and surface and of the extra-terrestrial solar irradiance. In this paper, the trending of the OMPS key calibration parameters in the past four years is shown. The impacts of these parameters on OMPS earth view radiance and albedo for nadir sensors are analyzed.
Ding Liang, Fuzhong Weng, Chunhui Pan, Ninghai Sun
IGARSS1
2016 Analysis of OMPS in-flight CCD dark current degradation
abstract
This paper presents the in-flight charge-coupled device (CCD) dark characterization of the Suomi National Polar-orbiting Partnership (S-NPP) Ozone Mapping Profiler Suite (OMPS). Data from OMPS's three different CCD detector arrays have been collected to characterize in-flight detector behaviors. It is of our primary interest in the evolution and trend of the dark current as well as the signal distribution beyond the first 4 years of the mission. Based on in-situ measurements obtained during the prelaunch calibration, we monitor changes in the CCD dark current on the pixel level in order to validate the in-flight OMPS dark calibration. The dark current change along with the Random Telegraph Signal (RTS) and the South Atlantic Anomaly (SAA) are studied while focusing on the influence of the Sensor Data Record's quality in terms of radiance errors. Our results show that current in-flight dark calibration provides reasonable Sensor Data Records (SDRs) and Environmental Data Records (EDRs) with an error of less than ∼0.1% on average in the earth-view radiance. However, calibration error inside of the SAA region of influence due to transients is sizable for wavelengths less than 302 nm where the signal-to-noise ratio is low.
Chunhui Pan, Fuzhong Weng, Trevor Beck, Ding Liang, Eve-Marie Devaliere, Wanchun Chen, Shuoguo Ding
IGARSS4
2014 Evaluation of the impact of a new quality control method on assimilation of CrIS data in HWRF-GSI
abstract
In this paper we evaluate the potential of assimilation of Cross-track Infrared Sounder (CrIS) radiance in the warm start HWRF system for Hurricane Sandy forecast using collocated Visible Infrared Imaging Radiometer Suite (VIIRS) cloud product for cloud detection and CrIS channels clearing. We examine the CrIS data bias correction and quality control procedure in GSI. Then we compare the cloud parameters retrieved from GSI stand-alone algorithm with those from CrIS/VIIRS collocated cloud products. And the impact of applying CrIS/VIIRS collocated cloud products on CrIS quality control in GSI has been evaluated.
Ding Liang, Fuzhong Weng
IGARSS1
2013 Evaluating parallel logistic regression models
abstract
Logistic regression (LR) has been widely used in applications of machine learning, thanks to its linear model. However, when the size of training data is very large, even such a linear model can consume excessive memory and computation time. To tackle both resource and computation scalability in a big-data setting, we evaluate and compare different approaches in distributed platform, parallel algorithm, and sublinear approximation. Our empirical study provides design guidelines for choosing the most effective combination for the performance requirement of a given application.
Haoruo Peng, Ding Liang, Cyrus Choi
IEEE BigData2
2012 Assessments of F18 special sensor microwave imager/sounder measurements for weather and climate applications
abstract
The Defense Meteorological Satellite Program's (DMSP) F-18 Special Sensor Microwave Imager/Sounder (SSMIS) brightness temperature difference (O-B) between observations and simulations from the Community Radiative Transfer Model (CRTM) based on 6 hour forecast fields of the National Centers for Environmental Prediction (NCEP) Global Forest System (GFS) are analyzed in this paper. It shows that after the current GSI bias correction, the O-B from SSMIS F-18 lower atmospheric sounding (LAS) channels are still significant and depend on nodes, seasons, latitudes and channels.
Ding Liang, Fuzhong Weng, Yong Chen 0011
IGARSS1
2009 Bistatic Reflection and Transmission of Electromagnetic Scattering by Rough Surfaces with Large Heights and Slopes
abstract
In this paper, we study the electromagnetic scattering properties of 2-D rough surface with large slope and large height. The ridges on the surface have heights of about 20cm. In microwave remote sensing of land, these heights are larger than wavelength. By using a tapered incident wave, the surface fields are solved by using numerical solution of Maxwell equations. Method of Moment (MOM) is used to solve the surface integral equations and rooftop basis function and Galerkin's method are used. The bistatic reflection and transmission are then calculated from the surface fields. Then the reflectivity from sastrugi over layered snow is calculated by solving multilayer radiative transfer (RT) equation with bistatic reflection and transmission coefficients as the boundary condition. We compare the electromagnetic scattering properties between Sastrugi rough surface and smooth surface. We show in this paper that for the sastrugi case, transmission angle can be larger than incident angle when incident from air to snow. This results in total internal reflection when the second layer of snow beneath sastrugi has a smaller permissivity and larger reflectivity than smooth surface.
Ding Liang, Peng Xu 0007, Kun-Shan Chen, Zhiqian Gui, Leung Tsang
IGARSS (2)1
2009 Comparison with CLPX II Airborne Data using DMRT Model
abstract
In this paper, we considered a physical-based model which use numerical solution of Maxwell Equations in three-dimensional simulations and apply into Dense Media Radiative Theory (DMRT). The model is validated in two specific dataset from the second Cold Land Processes Experiment (CLPX II) at Alaska and Colorado. The data were all obtain by the Ku-band (13.95 GHz) observations using airborne imaging polarimetric scatterometer (POLSCAT). Snow is a densely packed media. To take into account the collective scattering and incoherent scattering, analytical Quasi-Crystalline Approximation (QCA) and Numerical Maxwell Equation Method of 3-D simulation (NMM3D) are used to calculate the extinction coefficient and phase matrix. DMRT equations were solved by iterative solution up to 2ndorder for the case of small optical thickness and full multiple scattering solution by decomposing the diffuse intensities into Fourier series was used when optical thickness exceed unity. It was shown that the model predictions agree with the field experiment not only co-polarization but also cross-polarization. For Alaska region, the input snow structure data was obtain by the in situ ground observations, while for Colorado region, we combined the VIC model to get the snow profile.
Xiaolan Xu, Ding Liang, Konstantinos Andreadis, Leung Tsang, Edward G. Josberger
IGARSS (2)2
2008 Modeling Active Microwave Remote Sensing of Multilayer Dry Snow using Dense Media Radiative Transfer Theory
abstract
In this paper, we model the backscattering coefficients of multi-layer dry snowpacks, based on Dense Media Radiative Transfer theory (DMRT) with the Quasicrystalline Approximation (QCA). The DMRT model accounts for adhesive aggregate effects, which leads to dense media Mie scattering by using a Sticky particle model. The same set of DMRT equations are used for modeling both active and passive remote sensing. The model is validated by using the Cold-Land Processes Field Experiment CLPX ground based polarimetric scatterometry observation at local-scale observation site (LSOS) and airborne polarimetric Ku-band scatterometer (POLSCAT) data at Fool-Creek, Fraser. The snow density profiles are from ground observation and grain sizes are fitting parameters. It shows that the co-polarization simulations are in good agreement with the data, the cross-polarization simulations are around 2 dB lower than ground based observation and 5 dB lower than airborne observation. With the same set of multi-layer snowpack profile, the QCA/DMRT model matched co-polarization backscattering coefficients and all 4 channels of brightness temperature observations simultaneously at LSOS. The cross-polarization simulation can be improved by 3-dimensional numerical solutions of Maxwell equations (NMM3D). Study at Fool-Creek shows that NMM3D/DMRT simulations can match both co-polarization and cross-polarization observations simultaneously.
Ding Liang, Leung Tsang, Simon Yueh, Xiaolan Xu
IGARSS (3)1
2008 The Effects of Layers in Dry Snow on Its Passive Microwave Emissions Using Dense Media Radiative Transfer Theory Based on the Quasicrystalline Approximation (QCA/DMRT)
abstract
A model for the microwave emissions of multilayer dry snowpacks, based on dense media radiative transfer (DMRT) theory with the quasicrystalline approximation (QCA), provides more accurate results when compared to emissions determined by a homogeneous snowpack and other scattering models. The DMRT model accounts for adhesive aggregate effects, which leads to dense media Mie scattering by using a sticky particle model. With the multilayer model, we examined both the frequency and polarization dependence of brightness temperatures (Tb's) from representative snowpacks and compared them to results from a single-layer model and found that the multilayer model predicts higher polarization differences, twice as much, and weaker frequency dependence. We also studied the temporal evolution of Tb from multilayer snowpacks. The difference between Tb's at 18.7 and 36.5 GHz can be 5 K lower than the single-layer model prediction in this paper. By using the snowpack observations from the Cold Land Processes Field Experiment as input for both multi- and single-layer models, it shows that the multilayer Tb's are in better agreement with the data than the single-layer model. With one set of physical parameters, the multilayer QCA/DMRT model matched all four channels of Tb observations simultaneously, whereas the single-layer model could only reproduce vertically polarized Tb's. Also, the polarization difference and frequency dependence were accurately matched by the multilayer model using the same set of physical parameters. Hence, algorithms for the retrieval of snowpack depth or water equivalent should be based on multilayer scattering models to achieve greater accuracy.
Ding Liang, Xiaolan Xu, Leung Tsang, Konstantinos Andreadis, Edward G. Josberger
IEEE Trans. Geosci. Remote. Sens.1
2007 Modeling multi-layer effects in passive microwave remote sensing of dry snow using Dense Media Radiative Transfer Theory (DMRT) based on quasicrystalline approximation
abstract
The Dense Media Radiative Transfer theory (DMRT) of Quasicrystalline Approximation of Mie scattering by sticky particles is used to study the multiple scattering effects in layered snow in microwave remote sensing. Results are illustrated for various snow profile characteristics. Polarization differences and frequency dependences of multilayer snow model are significantly different from that of the single-layer snow model. Comparisons are also made with CLPX data using snow parameters as given by the VIC model.
Ding Liang, Xiaolan Xu, Leung Tsang, Konstantinos Andreadis, Edward G. Josberger
IGARSS1
2007 Modeling Active Microwave Remote Sensing of Snow Using Dense Media Radiative Transfer (DMRT) Theory With Multiple-Scattering Effects
abstract
Dense media radiative transfer (DMRT) theory is used to study the multiple-scattering effects in active microwave remote sensing. Simplified DMRT phase matrices are obtained in the 1-2 frame. The simplified expressions facilitate solutions of the DMRT equations and comparisons with other phase matrices. First-order, second-order, and full multiple-scattering solutions of the DMRT equations are obtained. To solve the DMRT equation, we decompose the diffuse intensities into Fourier series in the azimuthal direction. Each harmonic is solved by the eigen-quadrature approach. The model is applied to the active microwave remote sensing of terrestrial snow. Full multiple-scattering effects are important as the optical thickness for snow at frequencies above 10 GHz often exceed unity. The results are illustrated as a function of frequency, incidence angle, and snow depth. The results show that cross polarization for the case of densely packed spheres can be significant and can be merely 6 to 8 dB below copolarization. The magnitudes of the cross polarization are consistent with the experimental observations. The results show that the active 13.5-GHz backscattering coefficients still have significant sensitivity to snow thickness even for snow thickness exceeding 1 m
Leung Tsang, Ding Liang, Zhongxin Li, Donald W. Cline, Yunhua Tan
IEEE Trans. Geosci. Remote. Sens.3
2006 Modeling Active Microwave Remote Sensing of Snow using Dense Media Radiative Transfer (DMRT) Theory with Multiple Scattering Effects
abstract
Dense media radiative transfer theory (DMRT) is used to study the multiple scattering effects in active microwave remote sensing. To solve the dense media radiative transfer equation, we decompose the diffuse intensities into Fourier series in the azimuthal direction. Each harmonic is solved by the eigen-quadrature approach. The solution includes full multiple scattering effects within DMRT. Comparisons are made with the first order and the second order solutions. The model is applied to active microwave remote sensing of terrestrial snow. Full multiple scattering effects are important as the optical thickness for snow often exceed unity. The results are illustrated as a function of frequency, incidence angle and snow depth. The results show that cross polarization can be significant and can be only 6 to 8 dB below co-polarization, a result that is consistent with experimental observations. Also we note that even at snow depth of more than one meter, the active 13.5 GHz backscattering coefficients still have significant sensitivity to snow thickness.
Leung Tsang, Ding Liang, Zhongxin Li, Donald W. Cline
IGARSS3