Yonggang Qi

dblp:139/7002 · DBLP profile ↗
← Back
40ranked-venue papers
10as first author
25since 2021 · last 2026
0000-0001-8280-3541ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 27 · 7 first-author · 15 since 2021Artificial intelligence and machine learning · 20 · 5 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 SACG++: Complex Sketch Generation via Representation-Enhanced Scale-Adaptive Classifier Guidance
Ke Li 0004, Jijin Hu, Lan Yang 0014, Yonggang Qi, Yi-Zhe Song
Int. J. Comput. Vis.5
2026 Collaborative Model and Data Adaptation at Test Time
Chunyun Zhang, Fujun Yang, Chaoran Cui, Shuai Gong, Wenna Wang, Xue Lin 0003, Yonggang Qi, Lei Zhu 0002
IEEE Trans. Circuits Syst. Video Technol.7
2025 VersaGen: Unleashing Versatile Visual Control for Text-to-Image Synthesis
abstract
Despite the rapid advancements in text-to-image (T2I) synthesis, enabling precise visual control remains a significant challenge. Existing works attempted to incorporate multi-facet controls (text and sketch), aiming to enhance the creative control over generated images. However, our pilot study reveals that the expressive power of humans far surpasses the capabilities of current methods. Users desire a more versatile approach that can accommodate their diverse creative intents, ranging from controlling individual subjects to manipulating the entire scene composition. We present VersaGen, a generative AI agent that enables versatile visual control in T2I synthesis. VersaGen admits four types of visual controls: i) single visual subject; ii) multiple visual subjects; iii) scene background; iv) any combination of the three above or merely no control at all. We train an adaptor upon a frozen T2I model to accommodate the visual information into the text-dominated diffusion process. We introduce three optimization strategies during the inference phase of VersaGen to improve generation results and enhance user experience. Comprehensive experiments on COCO and Sketchy validate the effectiveness and flexibility of VersaGen, as evidenced by both qualitative and quantitative results.
Lan Yang 0014, Yonggang Qi, Honggang Zhang 0002, Kaiyue Pang, Ke Li 0004, Yi-Zhe Song
AAAI3
2025 SAUGE: Taming SAM for Uncertainty-Aligned Multi-Granularity Edge Detection
abstract
Edge labels are typically at various granularity levels owing to the varying preferences of annotators, thus handling the subjectivity of per-pixel labels has been a focal point for edge detection. Previous methods often employ a simple voting strategy to diminish such label uncertainty or impose a strong assumption of labels with a pre-defined distribution, e.g., Gaussian. In this work, we unveil that the segment anything model (SAM) provides strong prior knowledge to model the uncertainty in edge labels. Our key insight is that the intermediate SAM features inherently correspond to object edges at various granularities, which reflects different edge options due to uncertainty. Therefore, we attempt to align uncertainty with granularity by regressing intermediate SAM features from different layers to object edges at multi-granularity levels. In doing so, the model can fully and explicitly explore diverse ``uncertainties'' in a data-driven fashion. Specifically, we inject a lightweight module (~ 1.5% additional parameters) into the frozen SAM to progressively fuse and adapt its intermediate features to estimate edges from coarse to fine. It is crucial to normalize the granularity level of human edge labels to match their innate uncertainty. For this, we simply perform linear blending to the real edge labels at hand to create pseudo labels with varying granularities. Consequently, our uncertainty-aligned edge detector can flexibly produce edges at any desired granularity (including an optimal one). Thanks to SAM, our model uniquely demonstrates strong generalizability for cross-dataset edge detection. Extensive experimental results on BSDS500, Muticue and NYUDv2 validate our model's superiority.
Xing Liufu, Chaolei Tan, Xiaotong Lin 0002, Yonggang Qi, Jinxuan Li, Jianfang Hu
AAAI4
2025 AnimateSketches: Animate Sketches with Instance-Aware Mask
abstract
Sketch animation, the transformation of static sketches into dynamic experiences, is an essential tool for visual expression. Current methods rely on global-level optimization and neglect instance-level priors, resulting in optimization mismatch and overfitting across multiple sketches. To solve this problem, information about the attention map, representing different instances is needed to guide the optimization. In this work, we propose AnimateSketches, a novel optimization-based framework that focuses on animating complex vectorized sketches. We introduce prompt-guided instance-aware mask generation (PGIM), which leverages attention maps from a pretrained diffusion model to guide the optimization of individual sketches. In addition, we use mask-based score distillation sampling (MSDS) to maintain the integrity of untargeted sketches. Quantitative and qualitative evaluations show the superiority of our approach over baseline methods in terms of visual quality and instance-level prompt-guided correspondence.
Haoge Deng, Jijin Hu, Yonggang Qi
ICASSP4
2025 Autoregressive Video Generation without Vector Quantization
abstract
This paper presents a novel approach that enables autoregressive video generation with high efficiency. We propose to reformulate the video generation problem as a non-quantized autoregressive modeling of temporal frame-by-frame prediction and spatial set-by-set prediction. Unlike raster-scan prediction in prior autoregressive models or joint distribution modeling of fixed-length tokens in diffusion models, our approach maintains the causal property of GPT-style models for flexible in-context capabilities, while leveraging bidirectional modeling within individual frames for efficiency. With the proposed approach, we train a novel video autoregressive model without vector quantization, termed NOVA. Our results demonstrate that NOVA surpasses prior autoregressive video models in data efficiency, inference speed, visual fidelity, and video fluency, even with a much smaller model capacity, i.e., 0.6B parameters. NOVA also outperforms state-of-the-art image diffusion models in text-to-image generation tasks, with a significantly lower training cost. Additionally, NOVA generalizes well across extended video durations and enables diverse zero-shot applications in one unified model. Code and models are publicly available at https://github.com/baaivision/NOVA.
Haoge Deng, Haiwen Diao, Yufeng Cui, Huchuan Lu, Shiguang Shan, Yonggang Qi
ICLR8
2025 FantasyTalking: Realistic Talking Portrait Generation via Coherent Motion Synthesis
abstract
Creating a realistic animatable avatar from a single static portrait remains challenging. Existing approaches often struggle to capture subtle facial expressions, the associated global body movements, and the dynamic background. To address these limitations, we propose a novel framework that leverages a pretrained video diffusion Transformer model to generate high-fidelity, coherent talking portraits with controllable motion dynamics. At the core of our work is a dual-stage audio-visual alignment strategy. In the first stage, we employ a clip-level training scheme to establish coherent global motion by aligning audio-driven dynamics across the entire scene, including the reference portrait, contextual objects, and background. In the second stage, we refine lip movements at the frame level using a lip-tracing mask, ensuring precise synchronization with audio signals. To preserve identity without compromising motion flexibility, we replace the commonly used reference network with a lightweight cross-attention module that effectively maintains facial consistency throughout the video. Furthermore, we integrate a motion intensity modulation module that explicitly controls facial keypoints and body joint trajectories, enabling fine-grained manipulation of portrait movements beyond mere lip motion. Extensive experimental results show that our proposed approach achieves higher quality with better realism, coherence, motion intensity, and identity preservation. Our demo, code, models can be found on this page: https://fantasy-amap.github.io/fantasy-talking/.
Mengchao Wang, Yaqi Fan, Yonggang Qi, Mu Xu
ACM Multimedia6
2025 Precise Diffusion Inversion: Towards Novel Samples and Few-Step Models
abstract
The diffusion inversion problem seeks to recover the latent generative trajectory of a diffusion model given a real image. Faithful inversion is critical for ensuring consistency in diffusion-based image editing. Prior works formulate this task as a fixed-point problem and solve it using numerical methods. However, achieving both accuracy and efficiency remains challenging, especially for few-step models and novel samples. In this paper, we propose ***PreciseInv***, a general-purpose test-time optimization framework that enables fast and faithful inversion in as few as two inference steps. Unlike root-finding methods, we reformulate inversion as a learning problem and introduce a dynamic programming-inspired strategy to recursively estimate a parameterized sequence of noise embeddings. This design leverages the smoothness of the diffusion latent space for accurate gradient-based optimization and ensures memory efficiency via recursive subproblem construction. We further provide a theoretical analysis of ***PreciseInv***'s convergence and derive a provable upper bound on its reconstruction error. Extensive experiments on COCO 2017, DarkFace, and a stylized cartoon dataset show that ***PreciseInv*** achieves state-of-the-art performance in both reconstruction quality and inference speed. Improvements are especially notable for few-step models and under distribution shifts. Moreover, precise inversion yields substantial gains in editing consistency for text-driven image manipulation tasks. Code is available at: https://github.com/panda7777777/PreciseInv
Jing Zuo, Luoping Cui, Chuang Zhu, Yonggang Qi
NeurIPS4
2025 Fuse2Match: Training-Free Fusion of Flow, Diffusion, and Contrastive Models for Zero-Shot Semantic Matching
abstract
Recent work shows that features from Stable Diffusion (SD) and contrastively pretrained models like DINO can be directly used for zero-shot semantic correspondence via naive feature concatenation. In this paper, we explore the stronger potential of Stable Diffusion 3 (SD3), a rectified flow-based model with a multimodal transformer backbone (MM-DiT). We show that semantic signals in SD3 are scattered across multiple timesteps and transformer layers, and propose a multi-level fusion scheme to extract discriminative features. Moreover, we identify that naive fusion across models suffers from inconsistent distributions, thus leading to suboptimal performance. To address this, we propose a simple yet effective confidence-aware feature fusion strategy that re-weights each model’s contribution based on prediction confidence scores derived from their matching uncertainties. Notably, this fusion approach is not only training-free but also enables per-pixel adaptive integration of heterogeneous features. The resulting representation, Fuse2Match, significantly outperforms strong baselines on SPair-71k, PF-Pascal, and PSC6K, validating the benefit of combining SD3, SD, and DINO through our proposed confidence-aware feature fusion. Code is available at https://github.com/panda7777777/fuse2match
Jing Zuo, Yonggang Qi, Yi-Zhe Song
NeurIPS3
2025 FastTalker: Real-time audio-driven talking face generation with 3D Gaussian
Keliang Chen, Fang Cui, Mao Ni, Shaoying Wang, Junlin Che, Yonggang Qi, Fangwei Zhang, Gan Guo, Yunxia Huang
Image Vis. Comput.8
2024 Scale-Adaptive Diffusion Model for Complex Sketch Synthesis
abstract
While diffusion models have revolutionized generative AI, their application to human sketch generation, especially in the creation of complex yet concise and recognizable sketches, remains largely unexplored. Existing efforts have primarily focused on vector-based sketches, limiting their ability to handle intricate sketch data. This paper introduces an innovative extension of diffusion models to pixellevel sketch generation, addressing the challenge of dynamically optimizing the guidance scale for classifier-guided diffusion. Our approach achieves a delicate balance between recognizability and complexity in generated sketches through scale-adaptive classifier-guided diffusion models, a scaling indicator, and the concept of a residual sketch. We also propose a three-phase sampling strategy to enhance sketch diversity and quality. Experiments on the QuickDraw dataset showcase the potential of diffusion models to push the boundaries of sketch generation, particularly in complex scenarios unattainable by vector-based methods.
Jijin Hu, Ke Li 0004, Yonggang Qi, Yi-Zhe Song
ICLR3
2024 Sketch2Seg: Sketch-Based Image Segmentation with Pre-trained Diffusion Model
Haoge Deng, Yonggang Qi
ICPR (23)4
2024 Recalibration and Reprocessing of the Long-Term FY-3 MERSI Historical Data
abstract
The MEdium Resolution Spectral Imager (MERSI) onboard the Fengyun-3 (FY-3) series satellites can provide the long-term series data with favorable spectral and spatial resolution on the global scale since 2008. Such datasets are valuable for the studies of climate change. However, due to the lack of stable and reliable onboard calibration equipment and inconsistent in-orbit calibration methods, the MRESI historical data have poor long-term stability and unreliable accuracy, which affects the quantitative application of the data. This study reveals the overall status of the FY-3A/B/C MERSI-I historical data and proposes the recalibration methods for the reflective solar bands (RSBs) and thermal emission bands (TEBs). For the RSBs, by using FY-3A as the radiative transfer reference, an integrated transfer calibration method is developed for the calibration of FY-3B, which is then used to recalibrate FY-3C. The degradation tracking model of FY-3 MERSI-I is established first in the recalibration process by integrating multiple calibration methods. Then, based on the overlapping observations over the Libyan Desert, the linear consistency transfer model of the reference and target satellites is established, and the consistent correction coefficient between them is obtained. For the TEBs, a retrospective transfer recalibration scheme is proposed to achieve the reevaluation of the in-orbit radiometric calibration parameters based on intercalibration and to conduct the recalibration of historical data without permanent dependence on reference instruments. All the historical data of FY-A/B/C MERSI-I (from February 2008 to March 2017) are reprocessed with the same calibration method. The reprocessed datasets show remarkable improvements in calibration accuracy and stability compared with the operational datasets. The overall radiometric biases are found to be small and highly stable during the entire mission cycle of the instrument. The calibration biases of reprocessed data are less than 3% and 0.5 K for the RSBs and TEBs, respectively, much better than those of the operational datasets. There are also substantial improvements in the seasonal fluctuations and deviation discontinuities. This reprocessed long-term MERSI data with high intersensor consistency can provide valuable insights into global climate monitoring and model assessment.
Na Xu 0001, Xingwei He 0004, Xiuqing Hu, Hanlie Xu, Ronghua Wu, Ling Sun 0003, Lin Chen 0017, Yonggang Qi, Peng Zhang 0024
IEEE Trans. Geosci. Remote. Sens.9
2023 Zero-Shot Everything Sketch-Based Image Retrieval, and in Explainable Style
abstract
This paper studies the problem of zero-short sketch-based image retrieval (ZS-SBIR), however with two significant differentiators to prior art (i) we tackle all variants (inter-category, intracategory, and cross datasets) of ZS-SBIR with just one network (“everything”), and (ii) we would really like to understand how this sketch-photo matching operates (“explainable”). Our key innovation lies with the realization that such a cross-modal matching problem could be reduced to comparisons of groups of key local patches - akin to the seasoned “bag-of-words” paradigm. Just with this change, we are able to achieve both of the aforementioned goals, with the added benefit of no longer requiring external semantic knowledge. Technically, ours is a transformer-based cross-modal network, with three novel components (i) a self-attention module with a learnable tokenizer to produce visual tokens that correspond to the most informative local regions, (ii) a cross-attention module to compute local correspondences between the visual tokens across two modalities, and finally (iii) a kernel-based relation network to assemble local putative matches and produce an overall similarity metric for a sketch-photo pair. Experiments show ours indeed delivers superior performances across all ZS-SBIR settings. The all important explainable goal is elegantly achieved by visualizing cross-modal token correspondences, and for the first time, via sketch to photo synthesis by universal replacement of all matched photo patches. Code and model are available at https://github.com/buptLinfy/ZSE-SBIR.
Fengyin Lin, Mingkang Li 0003, Da Li 0001, Timothy M. Hospedales, Yi-Zhe Song, Yonggang Qi
CVPR6
2023 SketchKnitter: Vectorized Sketch Generation with Diffusion Models
Haoge Deng, Yonggang Qi, Da Li 0001, Yi-Zhe Song
ICLR3
2023 SketchScene: Scene Sketch To Image Generation With Diffusion Models
abstract
Sketch is an abstract visual representation that can be recovered as natural photographs in the human mind. Many researchers are drawn to work on translating abstract sketches to natural photographs. Since conventional sketch-to-image models are designed to generate images with a single object as the subject, generating scene image with multiple classes of objects is a tricky problem. To tackle this challenge, we propose the first scene sketch-to-image generation method based on diffusion models. Our model uses an encoder to summarize the contour and class features of the scene sketch into a latent variable, and a decoder to reconstruct scene images from it. In scene sketch-to-image generation tasks, our method outperforms the state-of-the-art methods. Experiments also show that our model beats other methods in zero-shot general sketch-to-image generation. It demonstrates our model’s potential for full-domain image generation.
Zhenbei Wu, Haoge Deng, Di Kong, Jie Yang 0023, Yonggang Qi
ICME6
2022 A Diffusion-ReFinement Model for Sketch-to-Point Modeling
Di Kong, Yonggang Qi
ACCV (7)3
2022 DiffSketching: Sketch Control Image Synthesis with Diffusion Models
Di Kong, Fengyin Lin, Yonggang Qi
BMVC4
2022 XPNet: Cross-Domain Prototypical Network for Zero-Shot Sketch-Based Image Retrieval
Mingkang Li 0003, Yonggang Qi
PRCV (1)2
2022 Detection-by-tracking of traffic signs in videos
Yanting Zhang 0001, Zijian Wang 0010, Ruoning Song, Cairong Yan, Yonggang Qi
Appl. Intell.5
2022 Generative Sketch Healing
Yonggang Qi, Guoyao Su, Jie Yang 0023, Kaiyue Pang, Yi-Zhe Song
Int. J. Comput. Vis.1
2021 PQA: Perceptual Question Answering
abstract
Perceptual organization remains one of the very few established theories on the human visual system. It underpinned many pre-deep seminal works on segmentation and detection, yet research has seen a rapid decline since the preferential shift to learning deep models. Of the limited attempts, most aimed at interpreting complex visual scenes using perceptual organizational rules. This has however been proven to be sub-optimal, since models were unable to effectively capture the visual complexity in real-world imagery. In this paper, we rejuvenate the study of perceptual organization, by advocating two positional changes: (i) we examine purposefully generated synthetic data, instead of complex real imagery, and (ii) we ask machines to synthesize novel perceptually-valid patterns, instead of explaining existing data. Our overall answer lies with the introduction of a novel visual challenge – the challenge of perceptual question answering (PQA). Upon observing example perceptual question-answer pairs, the goal for PQA is to solve similar questions by generating answers entirely from scratch (see Figure 1). Our first contribution is therefore the first dataset of perceptual question-answer pairs, each generated specifically for a particular Gestalt principle. We then borrow insights from human psychology to design an agent that casts perceptual organization as a self-attention problem, where a proposed grid-to-grid mapping network directly generates answer patterns from scratch. Experiments show our agent to outperform a selection of naive and strong baselines. A human study however indicates that ours uses astronomically more data to learn when compared to an average human, necessitating future research (with or without our dataset).
Yonggang Qi, Aneeshan Sain, Yi-Zhe Song
CVPR1
2021 SketchLattice: Latticed Representation for Sketch Manipulation
abstract
The key challenge in designing a sketch representation lies with handling the abstract and iconic nature of sketches. Existing work predominantly utilizes either, (i) a pixelative format that treats sketches as natural images employing off-the-shelf CNN-based networks, or (ii) an elaborately designed vector format that leverages the structural information of drawing orders using sequential RNN-based methods. While the pixelative format lacks intuitive exploitation of structural cues, sketches in vector format are absent in most cases limiting their practical usage. Hence, in this paper, we propose a lattice structured sketch representation that not only removes the bottleneck of requiring vector data but also preserves the structural cues that vector data provides. Essentially, sketch lattice is a set of points sampled from the pixelative format of the sketch using a lattice graph. We show that our lattice structure is particularly amenable to structural changes that largely benefits sketch abstraction modeling for generation tasks. Our lattice representation could be effectively encoded using a graph model, that uses significantly fewer model parameters (13.5 times lesser) than existing state-of-the-art. Extensive experiments demonstrate the effectiveness of sketch lattice for sketch manipulation, including sketch healing and image-to-sketch synthesis.
Yonggang Qi, Guoyao Su, Pinaki Nath Chowdhury, Mingkang Li 0003, Yi-Zhe Song
ICCV1
2021 Towards Practical Sketch-Based 3D Shape Generation: The Role of Professional Sketches
abstract
In this paper, for the first time, we investigate the problem of generating 3D shapes from professional 2D sketches via deep learning. We target sketches done by professional artists, as these sketches are likely to contain more details than the ones produced by novices, and thus the reconstruction from such sketches poses a higher demand on the level of detail in the reconstructed models. This is importantly different to previous work, where the training and testing was conducted on either synthetic sketches or sketches done by novices. Novices sketches often depict shapes that are physically unrealistic, while models trained with synthetic sketches could not cope with the level of abstraction and style found in real sketches. To address this problem, we collected the first large-scale dataset of professional sketches, where each sketch is paired with a reference 3D shape, with a total of 1,500 professional sketches collected across 500 3D shapes. The dataset is available at http://sketchx.ai/downloads/. We introduce two bespoke designs within a deep adversarial network to tackle the imprecision of human sketches and the unique figure/ground ambiguity problem inherent to sketch-based reconstruction. We show that existing 3D shapes generation methods designed for images fail to be naively applied to our problem, and demonstrate the effectiveness of our method both qualitatively and quantitatively.
Yonggang Qi, Yulia Gryaditskaya, Honggang Zhang 0002, Yi-Zhe Song
IEEE Trans. Circuits Syst. Video Technol.2
2021 Toward Fine-Grained Sketch-Based 3D Shape Retrieval
abstract
In this paper we study, for the first time, the problem of fine-grained sketch-based 3D shape retrieval. We advocate the use of sketches as a fine-grained input modality to retrieve 3D shapes at instance-level - e.g., given a sketch of a chair, we set out to retrieve a specific chair from a gallery of all chairs. Fine-grained sketch-based 3D shape retrieval (FG-SBSR) has not been possible till now due to a lack of datasets that exhibit one-to-one sketch-3D correspondences. The first key contribution of this paper is two new datasets, consisting a total of 4,680 sketch-3D pairings from two object categories. Even with the datasets, FG-SBSR is still highly challenging because (i) the inherent domain gap between 2D sketch and 3D shape is large, and (ii) retrieval needs to be conducted at the instance level instead of the coarse category level matching as in traditional SBSR. Thus, the second contribution of the paper is the first cross-modal deep embedding model for FG-SBSR, which specifically tackles the unique challenges presented by this new problem. Core to the deep embedding model is a novel cross-modal view attention module which automatically computes the optimal combination of 2D projections of a 3D shape given a query sketch.
Anran Qi, Yulia Gryaditskaya, Jifei Song, Yongxin Yang, Yonggang Qi, Timothy M. Hospedales, Tao Xiang 0002, Yi-Zhe Song
IEEE Trans. Image Process.5
2020 SketchHealer: A Graph-to-Sequence Network for Recreating Partial Human Sketches
Guoyao Su, Yonggang Qi, Kaiyue Pang, Jie Yang 0023, Yi-Zhe Song
BMVC2
2020 S3Net: Graph Representational Network For Sketch Recognition
abstract
Sketches are distinctly different to photos. They are highly abstract and exhibit a severe lack of visual cues. Prior works have therefore explored additional traits unique to sketches to help recognition such as stroke ordering. In this paper, we pioneer in studying the role of structure in sketches, for the task of sketch recognition. In particular, we propose a novel graph representation specifically designed for sketches, which follows the inherent hierarchical relationship (segment-stroke-sketch”) of sketching elements. By conforming to this hierarchy, we also introduce ajoint network that encapsulates both the structural and temporal traits of sketches for sketch recognition, termed S3Net.S3Netemploys a recurrent neural network (RNN) to extract segmentlevel features, followed by a graph convolutional network (GCN) to aggregate them into sketch-level features. The RNN first encodes temporal cues in sketches while its outputs are used as node embedding to construct a hierarchical sketch-graph. The GCN module then takes in this sketchgraph to produce a structure-aware embedding for sketches. Extensive experiments on the QuickDraw dataset, exhibit superior performance over state-of-the-arts, surpassing them by over 4%. Ablative studies further demonstrate the effectiveness of the proposed structural graph for both inter-class, and intra-class feature discrimination. Code is available at: https://github.com/yanglan0225/s3net;.
Lan Yang 0014, Aneeshan Sain, Linpeng Li, Yonggang Qi, Honggang Zhang 0002, Yi-Zhe Song
ICME4
2020 Improved Traffic Sign Detection In Videos Through Reasoning Effective RoI Proposals
abstract
Traffic sign detection is an important task in assisted safety and autonomous driving. It is important to continuously detect the traffic signs emerged on the road. Currently, most object detection methods make independent detections based on single images. When we apply these methods directly to a video clip to detect traffic signs without taking into account temporal correlations among adjacent frames, missed detections or incorrect detections can frequently occur due to motion blur, size change, partial occlusion, and/or bad pose. In this paper, we fully exploit the temporal consistency of traffic sign detection in videos. More specifically, we incorporate information of adjacent frames with high confidence scores to enhance the discovery of potential objects in the missed or incorrect detected frames by “recovering” the missed RoI proposals or by “improving” the incorrect RoI proposals with low confidence scores. Our method can be regarded as a “detection-by-tracking” strategy, which results in a more robust detection performance in videos.
Yanting Zhang 0001, Yonggang Qi, Jie Yang 0023, Jenq-Neng Hwang
ICME2
2020 Sketch Fewer to Recognize More by Learning a Co-Regularized Sparse Representation
abstract
Categorizing free-hand human sketches has profound implications in applications such as human computer interaction and image retrieval. The task is non-trivial due to the iconic nature of sketches, signified by large variances in both appearance and structure when compared with photographs. Despite recent advances made by deep learning methods, the requirement of a large training set is commonly imposed making them impractical for real-world applications where training sketches are cumbersome to obtain - sketches have to be hand-drawn one by one other than crawled freely on the Internet. In this work, we aim to delve further into the data scarcity problem of sketch-related research, by proposing a few-shot sketch classification framework. The model is based on a co-regularized embedding algorithm where common/shareable parts of learned human sketches are exploited, thereby can embed query sketch into a co-regularized sparse representation space for few-shot classification. A new dataset of 8,000 part-level annotated sketches of 100 categories is also proposed to facilitate future research. Experiment shows that our approach can achieve an 5-way one-shot classification accuracy of 85%, and 20-way one-shot at 51%.
Yonggang Qi, Yi-Zhe Song
IEEE Trans. Circuits Syst. Video Technol.1
2019 Fengyun-4A Meteorological Satellite Data Service System
abstract
This paper introduces the characteristics and the new technologies applied in Fengyun-4A(FY-4A) Data Service System. The system is an important part of Fengyun-4(FY-4) series satellite ground application system, and also is the new generation platform for meteorological satellite data sharing services. The system integrates several high-performance servers, high availability disk array, large-scale automated tape library as well as operation system, database, storage software and application software, all of which constitutes archive and service application clusters. The system adopts cloud-computing technology, big data technology, WEBGIS technology to establish a stable, reliable and secure operational satellite data archive platform, which not only provides data retrieval and download services, but also provides various auxiliary means and information assistance.
Yonggang Qi, Jiashen Zhang, Di Xian
PDCAT1
2019 Unpaired Image-to-Sketch Translation Network for Sketch Synthesis
abstract
Image-to-sketch translation is to learn the mapping between an image and a corresponding human drawn sketch. Machine can be trained to mimic the human drawing process using a training set of aligned image-sketch pairs. However, to collect such paired data is quite expensive or even unavailable for many cases since sketches exhibit various level of abstractness and drawing preferences. Hence we present an approach for learning an image-to-sketch translation network via unpaired examples. A translation network, which can translate the representation in image latent space to sketch domain, is trained in unsupervised setting. To prevent the problem of representation shifting in cross-domain translation, a novel cycle+ consistency loss is explored. Experimental results on sketch recognition and sketch-based image retrieval demonstrate the effectiveness of our approach.
Yue Zhang 0016, Guoyao Su, Yonggang Qi, Jie Yang 0023
VCIP3
2018 CTSD: A Dataset for Traffic Sign Recognition in Complex Real-World Images
abstract
Traffic sign recognition (TSR) is an indispensable component for vision-based system of self-driving car. Promising results have been achieved which especially benefit from the rapid development of deep neural networks recently. However, there are few works focusing on the algorithms’ performances towards different complex conditions, such as weather and viewpoint variations. In this paper, we propose a new real-world TSR dataset, which is a dataset with several fine-grained conditions fine labeled involving weather, light condition, occlusion, distance, color fading and camera angle. Detailed and unbiased comparison results are reported about the performances of several state-of-the-arts on our proposed and five public TSR datasets. Experimental results demonstrate that current arts for TSR are still far from satisfactory especially when it comes to complex real-world cases.
Yanting Zhang 0001, Yonggang Qi, Jun Liu 0014, Jie Yang 0023
VCIP3
2017 Image retrieval by dense caption reasoning
abstract
Humans tend to understand image scene by recognizing visual elements, then conjecturing and inferring based on them, hence are able to search relevant images. In this paper, we concern about the problem of complex image retrieval by reasoning image dense captions, which is similar to the way of human perception for searching images. Specifically, we transform the problem of complex image retrieval into a dense captioning and scene graph matching issue by using structured language descriptions for retrieval. Experimental results on a novel proposed large-scale content-based image retrieval dataset demonstrate the rationality and effectiveness of our method.
Xinru Wei, Yonggang Qi, Jun Liu 0014, Fang Liu 0026
VCIP2
2016 Sketch-based image retrieval via Siamese convolutional neural network
abstract
Sketch-based image retrieval (SBIR) is a challenging task due to the ambiguity inherent in sketches when compared with photos. In this paper, we propose a novel convolutional neural network based on Siamese network for SBIR. The main idea is to pull output feature vectors closer for input sketch-image pairs that are labeled as similar, and push them away if irrelevant. This is achieved by jointly tuning two convolutional neural networks which linked by one loss function. Experimental results on Flickr15K demonstrate that the proposed method offers a better performance when compared with several state-of-the-art approaches.
Yonggang Qi, Yi-Zhe Song, Honggang Zhang 0002, Jun Liu 0014
ICIP1
2016 Fengyun-3 Series Meteorological Satellite Data Archiving and Service System
abstract
This paper introduces the characteristics and the new technologies applied in Fengyun-3 Series Meteorological Satellite Data Archiving and Service System. The system integrates several high-performance servers, high availability disk array, large-scale automated tape library as well as operation system, database, storage software and application software, all of which constitutes archive and service application clusters. The application of configured dynamic load container technology and multi-server processing load balancing and parallel computing technology makes the cluster a highly scalable, available, efficient operational system to realize all levels of satellite data archiving, customization, retrieval, dynamic information release, operation monitoring, management, and other functions. Internet-based full dataset sharing, distributed file system, spatiotemporal integration of database storage, WebGIS global satellite image distribution, visual processing and display technology are adopted in the system to improve data sharing and service. Completion of the system meets the needs of operational meteorological satellite, achieving daily 4 TB data storing and daily 1 TB data download service. The system running in NSMC/CMA is the first one to provide internet-based full archive dataset (online, nearline and offline) sharing service in the area of Chinese civil remote sensing satellite data services.
Jianmei Qian, Di Xian, Yonggang Qi
IGARSS4
2015 Making better use of edges via perceptual grouping
abstract
We propose a perceptual grouping framework that organizes image edges into meaningful structures and demonstrate its usefulness on various computer vision tasks. Our grouper formulates edge grouping as a graph partition problem, where a learning to rank method is developed to encode probabilities of candidate edge pairs. In particular, RankSVM is employed for the first time to combine multiple Gestalt principles as cue for edge grouping. Afterwards, an edge grouping based object proposal measure is introduced that yields proposals comparable to state-of-the-art alternatives. We further show how human-like sketches can be generated from edge groupings and consequently used to deliver state-of-the-art sketch-based image retrieval performance. Last but not least, we tackle the problem of freehand human sketch segmentation by utilizing the proposed grouper to cluster strokes into semantic object parts.
Yonggang Qi, Yi-Zhe Song, Tao Xiang 0002, Honggang Zhang 0002, Timothy M. Hospedales, Yi Li 0004, Jun Guo 0002
CVPR1
2015 Im2Sketch: Sketch generation by unconflicted perceptual grouping
Yonggang Qi, Jun Guo 0002, Yi-Zhe Song, Tao Xiang 0002, Honggang Zhang 0002, Zheng-Hua Tan
Neurocomputing1
2013 Sketching by perceptual grouping
abstract
Sketch is used for rendering the visual world since prehistoric times, and has become ubiquitous nowadays with the increasing availability of touchscreens on portable devices. However, how to automatically map images to sketches, a problem that has profound implications on applications such as sketch-based image retrieval, still remains open. In this paper, we propose a novel method that draws a sketch automatically from a single natural image. Sketch extraction is posed within an unified contour grouping framework, where perceptual grouping is first used to form contour segment groups, followed by a group-based contour simplification method that generate the final sketches. In our experiment, for the first time we pose sketch evaluation as a sketch-based object recognition problem and the results validate the effectiveness of our system over the state-of-the-arts alternatives.
Yonggang Qi, Jun Guo 0002, Yi Li 0004, Honggang Zhang 0002, Tao Xiang 0002, Yi-Zhe Song
ICIP1
2013 Perceptual grouping via untangling Gestalt principles
abstract
Gestalt principles, a set of conjoining rules derived from human visual studies, have been known to play an important role in computer vision. Many applications such as image segmentation, contour grouping and scene understanding often rely on such rules to work. However, the problem of Gestalt confliction, i.e., the relative importance of each rule compared with another, remains unsolved. In this paper, we investigate the problem of perceptual grouping by quantifying the confliction among three commonly used rules: similarity, continuity and proximity. More specifically, we propose to quantify the importance of Gestalt rules by solving a learning to rank problem, and formulate a multi-label graph-cuts algorithm to group image primitives while taking into account the learned Gestalt confliction. Our experiment results confirm the existence of Gestalt confliction in perceptual grouping and demonstrate an improved performance when such a confliction is accounted for via the proposed grouping algorithm. Finally, a novel cross domain image classification method is proposed by exploiting perceptual grouping as representation.
Yonggang Qi, Jun Guo 0002, Yi Li 0004, Honggang Zhang 0002, Tao Xiang 0002, Yi-Zhe Song, Zheng-Hua Tan
VCIP1
2013 A multi-label classification approach for Facial Expression Recognition
abstract
Facial Expression Recognition (FER) techniques have already been adopted in numerous multimedia systems. Plenty of previous research assumes that each facial picture should be linked to only one of the predefined affective labels. Nevertheless, in practical applications, few of the expressions are exactly one of the predefined affective states. Therefore, to depict the facial expressions more accurately, this paper proposes a multi-label classification approach for FER and each facial expression would be labeled with one or multiple affective states. Meanwhile, by modeling the relationship between labels via Group Lasso regularization term, a maximum margin multi-label classifier is presented and the convex optimization formulation guarantees a global optimal solution. To evaluate the performance of our classifier, the JAFFE dataset is extended into a multi-label facial expression dataset by setting threshold to its continuous labels marked in the original dataset and the labeling results have shown that multiple labels can output a far more accurate description of facial expression. At the same time, the classification results have verified the superior performance of our algorithm.
Kaili Zhao, Honggang Zhang 0002, Mingzhi Dong, Jun Guo 0002, Yonggang Qi, Yi-Zhe Song
VCIP5