EDBT 2026 Demo / reviewers in the wild / expert
Qian Yu 0002
dblp:16/3790-2
· DBLP profile ↗
32ranked-venue papers
4as first author
27since 2021 · last 2026
0000-0002-0538-7940ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 2 first-author · 19 since 2021Artificial intelligence and machine learning · 19 · 4 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ELBO-T2IAlign: A Generic ELBO-Based Method for Calibrating Pixel-Level Text-Image Alignment in Diffusion ModelsabstractDiffusion models excel at image generation. Recent studies have shown that these models not only generate high-quality images but also encode text-image alignment information through attention maps or loss functions. This information is valuable for various downstream tasks, including segmentation, text-guided image editing, and compositional image generation. However, current methods heavily rely on the assumption of perfect text-image alignment in diffusion models, which is not the case. In this paper, we propose using zero-shot referring image segmentation as a proxy task to evaluate the pixel-level image and class-level text alignment of popular diffusion models. We conduct an in-depth analysis of pixel-text misalignment in diffusion models from the perspective of training data bias. We find that misalignment occurs in images with small-sized, occluded, or rare object classes. Therefore, we propose ELBO-T2IAlign-a simple yet effective method to calibrate pixel-text alignment in diffusion models based on the evidence lower bound (ELBO) of likelihood. ELBO-T2IAlign is training-free and generic: it requires no additional annotations, model retraining, or architectural modifications, and it can be directly applied to different diffusion backbones. Extensive experiments on zero-shot referring image segmentation, text-guided image editing, and compositional image generation verify that the proposed calibration improves pixel-text alignment across complementary downstream tasks. Qin Zhou 0003, Jing Zhang 0017, Qian Yu 0002, Lu Sheng, Dong Xu 0001 |
IEEE Trans. Image Process. | 6 |
| 2025 | TrackGo: A Flexible and Efficient Method for Controllable Video GenerationabstractRecent years have seen substantial progress in diffusion-based controllable video generation. However, achieving precise control in complex scenarios, including fine-grained object parts, sophisticated motion trajectories, and coherent background movement, remains a challenge. In this paper, we introduce *TrackGo*, a novel approach that leverages free-form masks and arrows for conditional video generation. This method offers users with a flexible and precise mechanism for manipulating video content. We also propose the *TrackAdapter* for control implementation, an efficient and lightweight adapter designed to be seamlessly integrated into the temporal self-attention layers of a pretrained video generation model. This design leverages our observation that the attention map of these layers can accurately activate regions corresponding to motion in videos. Our experimental results demonstrate that our new approach, enhanced by the TrackAdapter, achieves state-of-the-art performance on key metrics such as FVD, FID, and ObjMC scores. Haitao Zhou, Chuang Wang 0008, Jinlin Liu, Dongdong Yu, Qian Yu 0002, Changhu Wang |
AAAI | 6 |
| 2025 | Empowering LLMs to Understand and Generate Complex Vector GraphicsabstractThe unprecedented advancements in Large Language Models (LLMs) have profoundly impacted natural language processing but have yet to fully embrace the realm of scalable vector graphics (SVG) generation. While LLMs encode partial knowledge of SVG data from web pages during training, recent findings suggest that semantically ambiguous and tokenized representations within LLMs may result in hallucinations in vector primitive predictions. Additionally, LLM training typically lacks modeling and understanding of the rendering sequence of vector paths, which can lead to occlusion between output vector primitives. In this paper, we present LLM4SVG, an initial yet substantial step toward bridging this gap by enabling LLMs to better understand and generate vector graphics. LLM4SVG facilitates a deeper understanding of SVG components through learnable semantic tokens, which precisely encode these tokens and their corresponding properties to generate semantically aligned SVG outputs. Using a series of learnable semantic tokens, a structured dataset for instruction following is developed to support comprehension and generation across two primary tasks. Our method introduces a modular architecture to existing large language models, integrating semantic tags, vector instruction encoders, fine-tuned commands, and powerful LLMs to tightly combine geometric, appearance, and language information. To overcome the scarcity of SVG-text instruction data, we developed an automated data generation pipeline that collected our SVGX-SFT Dataset, consisting of high-quality human-designed SVGs and 580k SVG instruction following data specifically crafted for LLM training, which facilitated the adoption of the supervised fine-tuning strategy popular in LLM development. By exploring various training strategies, we developed LLM4SVG, which significantly moves beyond optimized rendering-based approaches and language-model-based baselines to achieve remarkable results in human evaluation tasks. Code, model, and data will be released at: https://ximinng.github.io/LLM4SVGProject/ Ximing Xing, Juncheng Hu 0001, Guotao Liang, Jing Zhang 0017, Dong Xu 0001, Qian Yu 0002 |
CVPR | 6 |
| 2025 | VectorPainter: Advanced Stylized Vector Graphics Synthesis Using Stroke-Style PriorsabstractWe introduce VectorPainter, a novel framework designed for reference-guided text-to-vector-graphics synthesis. Based on our observation that the style of strokes can be an important aspect to distinguish different artists, our method reforms the task into synthesizing a desired vector graphic by rearranging stylized strokes, which are vectorized from the reference images. Specifically, our method first converts the pixels of the reference image into a series of vector strokes, and then generates a vector graphic based on the input text description by optimizing the positions and colors of these vector strokes. To precisely capture the style of the reference image in the vectorized strokes, we propose an innovative vectorization method that employs an imitation learning strategy. To preserve the style of the strokes throughout the generation process, we introduce a style-preserving loss function. Extensive experiments have been conducted to demonstrate the superiority of our approach over existing works in stylized vector graphics synthesis, as well as the effectiveness of the various components of our method. Code, model, and data will be released at: https://hjc-owo.github.io/VectorPainterProject/ Juncheng Hu 0001, Ximing Xing, Jing Zhang 0017, Qian Yu 0002 |
ICME | 4 |
| 2025 | Multi-Object Sketch Animation with Grouping and Motion Trajectory Priors
Guotao Liang, Juncheng Hu 0001, Ximing Xing, Jing Zhang 0017, Qian Yu 0002 |
ACM Multimedia | 5 |
| 2025 | CAD-Coder: Text-to-CAD Generation with Chain-of-Thought and Geometric RewardabstractIn this work, we introduce CAD-Coder, a novel framework that reformulates text-to-CAD as the generation of CadQuery scripts—a Python-based, parametric CAD language.
This representation enables direct geometric validation, a richer modeling vocabulary, and seamless integration with existing LLMs.
To further enhance code validity and geometric fidelity, we propose a two-stage learning pipeline: (1) supervised fine-tuning on paired text–CadQuery data, and (2) reinforcement learning with Group Reward Policy Optimization (GRPO), guided by a CAD-specific reward comprising both a geometric reward (Chamfer Distance) and a format reward.
We also introduce a chain-of-thought (CoT) planning process to improve model reasoning, and construct a large-scale, high-quality dataset of 110K text–CadQuery–3D model triplets and 1.5K CoT samples via an automated pipeline. Extensive experiments demonstrate that CAD-Coder enables LLMs to generate diverse, valid, and complex CAD models directly from natural language, advancing the state of the art of text-to-CAD generation and geometric reasoning. Yandong Guan, Xilin Wang, Ximing Xing, Jing Zhang 0017, Dong Xu 0001, Qian Yu 0002 |
NeurIPS | 6 |
| 2025 | ViewCraft3D: High-fidelity and View-Consistent 3D Vector Graphics Synthesisabstract3D vector graphics play a crucial role in various applications including 3D shape retrieval, conceptual design, and virtual reality interactions due to their ability to capture essential structural information with minimal representation.
While recent approaches have shown promise in generating 3D vector graphics, they often suffer from lengthy processing times and struggle to maintain view consistency.
To address these limitations, we propose VC3D (**V**iew**C**raft**3D**), an efficient method that leverages 3D priors to generate 3D vector graphics.
Specifically, our approach begins with 3D object analysis, employs a geometric extraction algorithm to fit 3D vector graphics to the underlying structure, and applies view-consistent refinement process to enhance visual quality.
Our comprehensive experiments demonstrate that VC3D outperforms previous methods in both qualitative and quantitative evaluations, while significantly reducing computational overhead. The resulting 3D sketches maintain view consistency and effectively capture the essential characteristics of the original objects. Chuang Wang 0008, Haitao Zhou, Qian Yu 0002 |
NeurIPS | 4 |
| 2025 | SVGDreamer++: Advancing Editability and Diversity in Text-Guided SVG GenerationabstractRecently, text-guided scalable vector graphics (SVG) synthesis has shown great promise in domains like iconography and sketching. However, existing Text-to-SVG methods often face challenges in editability, visual quality, and diversity. To address these issues, we propose a novel framework for text-guided SVG synthesis that significantly enhances editability, quality, and diversity. To enhance the editability of output SVGs, we introduce a Hierarchical Image VEctorization (HIVE) framework that operates at the semantic object level and supervises the optimization of components within the vector object. This approach facilitates the decoupling of vector graphics into distinct objects and component levels. Our proposed HIVE algorithm, informed by image segmentation priors, not only ensures a more precise representation of vector graphics but also enables fine-grained editing capabilities within vector objects. To improve the diversity of output SVGs, we present a Vectorized Particle-based Score Distillation (VPSD) approach. VPSD addresses over-saturation issues in existing methods and enhances sample diversity. A pre-trained reward model is incorporated to re-weight vector particles, improving aesthetic appeal and enabling faster convergence. Additionally, we design a novel adaptive vector primitives control strategy, which allows for the dynamic adjustment of the number of primitives, thereby enhancing the presentation of graphic details. Extensive experiments validate the effectiveness of the proposed method, demonstrating its superiority over baseline methods in terms of editability, visual quality, and diversity. We also show that our new method supports up to six distinct vector styles, capable of generating high-quality vector assets suitable for stylized vector design and poster design. Ximing Xing, Qian Yu 0002, Chuang Wang 0008, Haitao Zhou, Jing Zhang 0017, Dong Xu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Diffusion Model is Secretly a Training-Free Open Vocabulary Semantic SegmenterabstractThe pre-trained text-image discriminative models, such as CLIP, has been explored for open-vocabulary semantic segmentation with unsatisfactory results due to the loss of crucial localization information and awareness of object shapes. Recently, there has been a growing interest in expanding the application of generative models from generation tasks to semantic segmentation. These approaches utilize generative models either for generating annotated data or extracting features to facilitate semantic segmentation. This typically involves generating a considerable amount of synthetic data or requiring additional mask annotations. To this end, we uncover the potential of generative text-to-image diffusion models (e.g., Stable Diffusion) as highly efficient open-vocabulary semantic segmenters, and introduce a novel training-free approach named DiffSegmenter. The insight is that to generate realistic objects that are semantically faithful to the input text, both the complete object shapes and the corresponding semantics are implicitly learned by diffusion models. We discover that the object shapes are characterized by the self-attention maps while the semantics are indicated through the cross-attention maps produced by the denoising U-Net, forming the basis of our segmentation results. Additionally, we carefully design effective textual prompts and a category filtering mechanism to further enhance the segmentation results. Extensive experiments on three benchmark datasets show that the proposed DiffSegmenter achieves impressive results for open-vocabulary semantic segmentation. Xiawei Li, Jing Zhang 0017, Qingyuan Xu, Qin Zhou 0003, Qian Yu 0002, Lu Sheng, Dong Xu 0001 |
IEEE Trans. Image Process. | 6 |
| 2024 | Multi-Modality Affinity Inference for Weakly Supervised 3D Semantic Segmentationabstract3D point cloud semantic segmentation has a wide range of applications. Recently, weakly supervised point cloud segmentation methods have been proposed, aiming to alleviate the expensive and laborious manual annotation process by leveraging scene-level labels. However, these methods have not effectively exploited the rich geometric information (such as shape and scale) and appearance information (such as color and texture) present in RGB-D scans. Furthermore, current approaches fail to fully leverage the point affinity that can be inferred from the feature extraction network, which is crucial for learning from weak scene-level labels. Additionally, previous work overlooks the detrimental effects of the long-tailed distribution of point cloud data in weakly supervised 3D semantic segmentation. To this end, this paper proposes a simple yet effective scene-level weakly supervised point cloud segmentation method with a newly introduced multi-modality point affinity inference module. The point affinity proposed in this paper is characterized by features from multiple modalities (e.g., point cloud and RGB), and is further refined by normalizing the classifier weights to alleviate the detrimental effects of long-tailed distribution without the need of the prior of category distribution. Extensive experiments on the ScanNet and S3DIS benchmarks verify the effectiveness of our proposed method, which outperforms the state-of-the-art by ~4% to ~ 6% mIoU. Codes are released at https://github.com/Sunny599/AAAI24-3DWSSG-MMA. Xiawei Li, Qingyuan Xu, Jing Zhang 0017, Qian Yu 0002, Lu Sheng, Dong Xu 0001 |
AAAI | 5 |
| 2024 | Data-Free Generalized Zero-Shot LearningabstractDeep learning models have the ability to extract rich knowledge from large-scale datasets. However, the sharing of data has become increasingly challenging due to concerns regarding data copyright and privacy. Consequently, this hampers the effective transfer of knowledge from existing data to novel downstream tasks and concepts. Zero-shot learning (ZSL) approaches aim to recognize new classes by transferring semantic knowledge learned from base classes. However, traditional generative ZSL methods often require access to real images from base classes and rely on manually annotated attributes, which presents challenges in terms of data restrictions and model scalability. To this end, this paper tackles a challenging and practical problem dubbed as data-free zero-shot learning (DFZSL), where only the CLIP-based base classes data pre-trained classifier is available for zero-shot classification. Specifically, we propose a generic framework for DFZSL, which consists of three main components. Firstly, to recover the virtual features of the base data, we model the CLIP features of base class images as samples from a von Mises-Fisher (vMF) distribution based on the pre-trained classifier. Secondly, we leverage the text features of CLIP as low-cost semantic information and propose a feature-language prompt tuning (FLPT) method to further align the virtual image features and textual features. Thirdly, we train a conditional generative model using the well-aligned virtual image features and corresponding semantic text features, enabling the generation of new classes features and achieve better zero-shot generalization. Our framework has been evaluated on five commonly used benchmarks for generalized ZSL, as well as 11 benchmarks for the base-to-new ZSL. The results demonstrate the superiority and effectiveness of our approach. Our code is available in https://github.com/ylong4/DFZSL. Jing Zhang 0017, Qian Yu 0002, Lu Sheng, Dong Xu 0001 |
AAAI | 4 |
| 2024 | SVGDreamer: Text Guided SVG Generation with Diffusion ModelabstractRecently, text-guided scalable vector graphics (SVGs) synthesis has shown promise in domains such as iconography and sketch. However, existing text-to-SVG generation methods lack editability and struggle with visual quality and result diversity. To address these limitations, we propose a novel text-guided vector graphics synthesis method called SVGDreamer. SVGDreamer incorporates a semantic-driven image vectorization (SIVE) process that enables the decomposition of synthesis into foreground objects and background, thereby enhancing editability. Specifically, the SIVE process introduces attention-based primitive control and an attention-mask loss function for effective control and manipulation of individual elements. Additionally, we propose a Vectorized Particle-based Score Distillation (VPSD) approach to address issues of shape over-smoothing, color over-saturation, limited diversity, and slow convergence of the existing text-to-SVG generation methods by modeling SVGs as distributions of control points and colors. Furthermore, VPSD leverages a reward model to re-weight vector particles, which improves aesthetic appeal and accelerates convergence. Extensive experiments are conducted to validate the effectiveness of SVGDreamer, demonstrating its superiority over baseline methods in terms of editability, visual quality, and diversity. Project page: https://ximinng.github.io/SVGDreamer-project/ Ximing Xing, Haitao Zhou, Chuang Wang 0008, Jing Zhang 0017, Dong Xu 0001, Qian Yu 0002 |
CVPR | 6 |
| 2024 | 3D Reconstruction From a Single Sketch via View-Dependent Depth SamplingabstractReconstructing a 3D shape based on a single sketch image is challenging due to the inherent sparsity and ambiguity present in sketches. Existing methods lose fine details when extracting features to predict 3D objects from sketches. Upon analyzing the 3D-to-2D projection process, we observe that the density map, characterizing the distribution of 2D point clouds, can serve as a proxy to facilitate the reconstruction process. In this work, we propose a novel sketch-based 3D reconstruction model named SketchSampler. It initiates the process by translating a sketch through an image translation network into a more informative 2D representation, which is then used to generate a density map. Subsequently, a two-stage probabilistic sampling process is employed to reconstruct a 3D point cloud: first, recovering the 2D points (i.e., the x and y coordinates) by sampling the density map; and second, predicting the depth (i.e., the z coordinate) by sampling the depth values along the ray determined by each 2D point. Additionally, we convert the reconstructed point cloud into a 3D mesh for wider applications. To reduce ambiguity, we incorporate hidden lines in sketches. Experimental results demonstrate that our proposed approach significantly outperforms other baseline methods. Chenjian Gao, Xilin Wang, Qian Yu 0002, Lu Sheng, Jing Zhang 0017, Xiaoguang Han 0001, Yi-Zhe Song, Dong Xu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | PolarFormer: A Transformer-Based Method for Multi-Lesion Segmentation in Intravascular OCTabstractSeveral deep learning-based methods have been proposed to extract vulnerable plaques of a single class from intravascular optical coherence tomography (OCT) images. However, further research is limited by the lack of publicly available large-scale intravascular OCT datasets with multi-class vulnerable plaque annotations. Additionally, multi-class vulnerable plaque segmentation is extremely challenging due to the irregular distribution of plaques, their unique geometric shapes, and fuzzy boundaries. Existing methods have not adequately addressed the geometric features and spatial prior information of vulnerable plaques. To address these issues, we collected a dataset containing 70 pullback data and developed a multi-class vulnerable plaque segmentation model, called PolarFormer, that incorporates the prior knowledge of vulnerable plaques in spatial distribution. The key module of our proposed model is Polar Attention, which models the spatial relationship of vulnerable plaques in the radial direction. Extensive experiments conducted on the new dataset demonstrate that our proposed method outperforms other baseline methods. Code and data can be accessed via this link: https://github.com/sunjingyi0415/IVOCT-segementaion. Zhili Huang, Yifan Shao, Qiyong Li, Jinsong Li 0003, Qian Yu 0002 |
IEEE Trans. Medical Imaging | 8 |
| 2023 | Feature Decomposition for Reducing Negative Transfer: A Novel Multi-Task Learning Method for Recommender System (Student Abstract)abstractWe propose a novel multi-task learning method termed Feature Decomposition Network (FDN). The key idea of the proposed FDN is to reduce the phenomenon of feature redundancy by explicitly decomposing features into task-specific features and task-shared features with carefully designed constraints. Experimental results show that our proposed FDN can outperform the state-of-the-art (SOTA) methods by a noticeable margin on Ali-CCP. Jie Zhou 0029, Qian Yu 0002, Chuan Luo 0002, Jing Zhang 0017 |
AAAI | 2 |
| 2023 | HiNet: Novel Multi-Scenario & Multi-Task Learning with Hierarchical Information ExtractionabstractMulti-scenario & multi-task learning has been widely applied to many recommendation systems in industrial applications, wherein an effective and practical approach is to carry out multi-scenario transfer learning on the basis of the Mixture-of-Expert (MoE) architecture. However, the MoE-based method, which aims to project all information in the same feature space, cannot effectively deal with the complex relationships inherent among various scenarios and tasks, resulting in unsatisfactory performance. To tackle the problem, we propose a Hierarchical information extraction Network (HiNet) for multi-scenario and multi-task recommendation, which achieves hierarchical extraction based on coarse-to-fine knowledge transfer scheme. The multiple extraction layers of the hierarchical network enable the model to enhance the capability of transferring valuable information across scenarios while preserving specific features of scenarios and tasks. Furthermore, a novel scenario-aware attentive network module is proposed to model correlations between scenarios explicitly. Comprehensive experiments conducted on real-world industrial datasets from Meituan Meishi platform demonstrate that HiNet achieves a new state-of-the-art performance and significantly outperforms existing solutions. HiNet is currently fully deployed in two scenarios and has achieved 2.87% and 1.75% order quantity gain respectively. Jie Zhou 0029, Xianshuai Cao, Lin Bo, Chuan Luo 0002, Qian Yu 0002 |
ICDE | 7 |
| 2023 | Vision Transformer Based Multi-class Lesion Detection in IVOCT
Yifan Shao, Zhili Huang, Qiyong Li, Jinsong Li 0003, Qian Yu 0002 |
MICCAI (6) | 8 |
| 2023 | Distortion-aware Transformer in 360° Salient Object DetectionabstractWith the emergence of VR and AR, 360° data attracts increasing attention from the computer vision and multimedia communities. Typically, 360° data is projected into 2D ERP (equirectangular projection) images for feature extraction. However, existing methods cannot handle the distortions that result from the projection, hindering the development of 360-data-based tasks. Therefore, in this paper, we propose a Transformer-based model called DATFormer to address the distortion problem. We tackle this issue from two perspectives. Firstly, we introduce two distortion-adaptive modules. The first is a Distortion Mapping Module, which guides the model to pre-adapt to distorted features globally. The second module is a Distortion-Adaptive Attention Block that reduces local distortions on multi-scale features. Secondly, to exploit the unique characteristics of 360° data, we present a learnable relation matrix and use it as part of the positional embedding to further improve performance. Extensive experiments are conducted on three public datasets, and the results show that our model outperforms existing 2D SOD (salient object detection) and 360 SOD methods. The source code is available at https://github.com/yjzhao19981027/DATFormer/. Yinjie Zhao, Lichen Zhao, Qian Yu 0002, Lu Sheng, Jing Zhang 0017, Dong Xu 0001 |
ACM Multimedia | 3 |
| 2023 | DiffSketcher: Text Guided Vector Sketch Synthesis through Latent Diffusion ModelsabstractEven though trained mainly on images, we discover that pretrained diffusion models show impressive power in guiding sketch synthesis. In this paper, we present DiffSketcher, an innovative algorithm that creates \textit{vectorized} free-hand sketches using natural language input. DiffSketcher is developed based on a pre-trained text-to-image diffusion model. It performs the task by directly optimizing a set of Bézier curves with an extended version of the score distillation sampling (SDS) loss, which allows us to use a raster-level diffusion model as a prior for optimizing a parametric vectorized sketch generator. Furthermore, we explore attention maps embedded in the diffusion model for effective stroke initialization to speed up the generation process. The generated sketches demonstrate multiple levels of abstraction while maintaining recognizability, underlying structure, and essential visual details of the subject drawn. Our experiments show that DiffSketcher achieves greater quality than prior work. The code and demo of DiffSketcher can be found at https://ximinng.github.io/DiffSketcher-project/. Ximing Xing, Chuang Wang 0008, Haitao Zhou, Jing Zhang 0017, Qian Yu 0002, Dong Xu 0001 |
NeurIPS | 5 |
| 2023 | SketchInverter: Multi-Class Sketch-Based Image Generation via GAN InversionabstractThis paper proposes the first GAN inversion-based method for multi-class sketch-based image generation (MCSBIG). MC-SBIG is a challenging task that requires strong prior knowledge due to the significant domain gap between sketches and natural images. Existing learning-based approaches rely on a large-scale paired dataset to learn the mapping between these two image modalities. However, since the public paired sketch-photo data are scarce, it is struggling for learning-based methods to achieve satisfactory results. In this work, we introduce a new approach based on GAN inversion, which can utilize a powerful pretrained generator to facilitate image generation from a given sketch. Our GAN inversion-based method has two advantages: 1. it can freely take advantage of the prior knowledge of a pretrained image generator; 2. it allows the proposed model to focus on learning the mapping from a sketch to a low-dimension latent code, which is a much easier task than directly mapping to a high-dimension natural image. We also present a novel shape loss to improve generation quality further. Extensive experiments are conducted to show that our method can produce sketch-faithful and photo-realistic images and significantly outperform the baseline methods. Zirui An, Jingbo Yu, Runtao Liu, Chuang Wang 0008, Qian Yu 0002 |
WACV | 5 |
| 2023 | GA-Sketching: Shape Modeling from Multi-View Sketching with Geometry-Aligned Deep Implicit FunctionsabstractAbstract Sketch‐based shape modeling aims to bridge the gap between 2D drawing and 3D modeling by providing an intuitive and accessible approach to create 3D shapes from 2D sketches. However, existing methods still suffer from limitations in reconstruction quality and multi‐view interaction friendliness, hindering their practical application. This paper proposes a faithful and user‐friendly iterative solution to tackle these limitations by learning geometry‐aligned deep implicit functions from one or multiple sketches. Our method lifts 2D sketches to volume‐based feature tensors, which align strongly with the output 3D shape, enabling accurate reconstruction and faithful editing. Such a geometry‐aligned feature encoding technique is well‐suited to iterative modeling since features from different viewpoints can be easily memorized or aggregated. Based on these advantages, we design a unified interactive system for sketch‐based shape modeling. It enables users to generate the desired geometry iteratively by drawing sketches from any number of viewpoints. In addition, it allows users to edit the generated surface by making a few local modifications. We demonstrate the effectiveness and practicality of our method with extensive experiments and user studies, where we found that our method outperformed existing methods in terms of accuracy, efficiency, and user satisfaction. The source code of this project is available at https://github.com/LordLiang/GA‐Sketching . Jie Zhou 0029, Zhongjin Luo, Qian Yu 0002, Xiaoguang Han 0001, Hongbo Fu 0001 |
Comput. Graph. Forum | 3 |
| 2023 | Slow Motion Matters: A Slow Motion Enhanced Network for Weakly Supervised Temporal Action LocalizationabstractWeakly supervised temporal action localization (WTAL) aims to localize actions in untrimmed videos with only weak supervision information (e.g., video-level labels). Most existing models handle all input videos with a fixed temporal scale. However, such models are not sensitive to actions whose pace of the movements is different from the “normal” speed, especially slow-motion action instances, which complete the movements with a much slower speed than their counterparts with a “normal” speed. Here arises the slow-motion blurred issue: It is hard to explore salient slow-motion information from videos at normal speed. In this paper, we propose a novel framework termed Slow Motion Enhanced Network (SMEN) to improve the ability of a WTAL network by compensating its sensitivity on slow-motion action segments. The proposed SMEN comprises a Mining module and a Localization module. The mining module generates mask to mine slow-motion-related features by utilizing the relationships between the normal motion and slow motion; while the localization module leverages the mined slow-motion features as complementary information to improve the temporal action localization results. Our proposed framework can be easily adapted by existing WTAL networks and enable them be more sensitive to slow-motion actions. Extensive experiments on three benchmarks are conducted, which demonstrate the high performance of our proposed framework. Qian Yu 0002, Dong Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | MedoidsFormer: A Strong 3D Object Detection Backbone by Exploiting Interaction With Adjacent Medoid TokensabstractIn this paper, we propose MedoidsFormer, a novel transformer-based backbone equipped with a self-attention mechanism that is tailored explicitly to LiDAR-based 3D object detection. Unlike 2D object detection, the proportion of target objects to the input scene is much smaller, and their distribution is significantly sparser in 3D object detection. Given these observations, we introduce a new self-attention mechanism called Medoids Attention, focusing on exploiting interactions within surrounding regions, which not only reduces computation and memory costs but obtains discriminative context information. Instead of aggregating tokens from adjacent areas, we present a dynamic semantic-aware token mining process through k-Medoids clustering to direct select representative tokens for attention modeling. Our proposed method shows consistent improvement over existing 3D object detectors through extensive experiments and achieves state-of-the-art performance on the large-scale Waymo Open Dataset. We also conduct comprehensive ablation studies to verify the efficacy of the new self-attention mechanism and provide thorough insights. Xiaoyu Tian, Qian Yu 0002, Jun-Hai Yong, Dong Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | SketchSampler: Sketch-Based 3D Reconstruction via View-Dependent Depth Sampling
Chenjian Gao, Qian Yu 0002, Lu Sheng, Yi-Zhe Song, Dong Xu 0001 |
ECCV (1) | 2 |
| 2022 | Human-Centric Spatio-Temporal Video Grounding With Visual TransformersabstractIn this work, we introduce a novel task – Human-centric Spatio-Temporal Video Grounding (HC-STVG). Unlike the existing referring expression tasks in images or videos, by focusing on humans, HC-STVG aims to localize a spatio-temporal tube of the target person from an untrimmed video based on a given textural description. This task is useful, especially for healthcare and security related applications, where the surveillance videos can be extremely long but only a specific person during a specific period is concerned. HC-STVG is a video grounding task that requires both spatial (where) and temporal (when) localization. Unfortunately, the existing grounding methods cannot handle this task well. We tackle this task by proposing an effective baseline method named Spatio-Temporal Grounding with Visual Transformers (STGVT), which utilizes Visual Transformers to extract cross-modal representations for video-sentence matching and temporal localization. To facilitate this task, we also contribute an HC-STVG datasetThe new dataset is available athttps://github.com/tzhhhh123/HC-STVG. consisting of 5,660 video-sentence pairs on complex multi-person scenes. Specifically, each video lasts for 20 seconds, pairing with a natural query sentence with an average of 17.25 words. Extensive experiments are conducted on this dataset, demonstrating that the newly-proposed method outperforms the existing baseline methods. Zongheng Tang, Yue Liao, Si Liu 0001, Guanbin Li, Xiaojie Jin 0004, Qian Yu 0002, Dong Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2021 | STVGBert: A Visual-linguistic Transformer based Framework for Spatio-temporal Video GroundingabstractSpatio-temporal video grounding (STVG) aims to localize a spatio-temporal tube of a target object in an untrimmed video based on a query sentence. In this work, we propose a one-stage visual-linguistic transformer based framework called STVGBert for the STVG task, which can simultaneously localize the target object in both spatial and temporal domains. Specifically, without resorting to pregenerated object proposals, our STVGBert directly takes a video and a query sentence as the input, and then produces the cross-modal features by using the newly introduced cross-modal feature learning module ST-ViLBert. Based on the cross-modal features, our method then generates bounding boxes and predicts the starting and ending frames to produce the predicted object tube. To the best of our knowledge, our STVGBert is the first one-stage method, which can handle the STVG task without relying on any pre-trained object detectors. Comprehensive experiments demonstrate our newly proposed framework outperforms the state-of-the-art multi-stage methods on two benchmark datasets Vid-STG and HC-STVG. Qian Yu 0002, Dong Xu 0001 |
ICCV | 2 |
| 2021 | Fine-Grained Instance-Level Sketch-Based Image Retrieval
Qian Yu 0002, Jifei Song, Yi-Zhe Song, Tao Xiang 0002, Timothy M. Hospedales |
Int. J. Comput. Vis. | 1 |
| 2018 | SketchyScene: Richly-Annotated Scene Sketches
Changqing Zou, Qian Yu 0002, Ruofei Du, Haoran Mo, Yi-Zhe Song, Tao Xiang 0002, Chengying Gao, Baoquan Chen, Hao (Richard) Zhang |
ECCV (15) | 2 |
| 2017 | Deep Spatial-Semantic Attention for Fine-Grained Sketch-Based Image RetrievalabstractHuman sketches are unique in being able to capture both the spatial topology of a visual object, as well as its subtle appearance details. Fine-grained sketch-based image retrieval (FG-SBIR) importantly leverages on such fine-grained characteristics of sketches to conduct instance-level retrieval of photos. Nevertheless, human sketches are often highly abstract and iconic, resulting in severe misalignments with candidate photos which in turn make subtle visual detail matching difficult. Existing FG-SBIR approaches focus only on coarse holistic matching via deep cross-domain representation learning, yet ignore explicitly accounting for fine-grained details and their spatial context. In this paper, a novel deep FG-SBIR model is proposed which differs significantly from the existing models in that: (1) It is spatially aware, achieved by introducing an attention module that is sensitive to the spatial position of visual details: (2) It combines coarse and fine semantic information via a shortcut connection fusion block: and (3) It models feature correlation and is robust to misalignments between the extracted features across the two domains by introducing a novel higher-order learnable energy function (HOLEF) based loss. Extensive experiments show that the proposed deep spatial-semantic attention model significantly outperforms the state-of-the-art. Jifei Song, Qian Yu 0002, Yi-Zhe Song, Tao Xiang 0002, Timothy M. Hospedales |
ICCV | 2 |
| 2017 | Sketch-a-Net: A Deep Neural Network that Beats Humans
Qian Yu 0002, Yongxin Yang, Feng Liu 0036, Yi-Zhe Song, Tao Xiang 0002, Timothy M. Hospedales |
Int. J. Comput. Vis. | 1 |
| 2016 | Sketch Me That ShoeabstractWe investigate the problem of fine-grained sketch-based image retrieval (SBIR), where free-hand human sketches are used as queries to perform instance-level retrieval of images. This is an extremely challenging task because (i) visual comparisons not only need to be fine-grained but also executed cross-domain, (ii) free-hand (finger) sketches are highly abstract, making fine-grained matching harder, and most importantly (iii) annotated cross-domain sketch-photo datasets required for training are scarce, challenging many state-of-the-art machine learning techniques. In this paper, for the first time, we address all these challenges, providing a step towards the capabilities that would underpin a commercial sketch-based image retrieval application. We introduce a new database of 1,432 sketchphoto pairs from two categories with 32,000 fine-grained triplet ranking annotations. We then develop a deep tripletranking model for instance-level SBIR with a novel data augmentation and staged pre-training strategy to alleviate the issue of insufficient fine-grained training data. Extensive experiments are carried out to contribute a variety of insights into the challenges of data sufficiency and over-fitting avoidance when training deep networks for finegrained cross-domain ranking tasks. Qian Yu 0002, Feng Liu 0036, Yi-Zhe Song, Tao Xiang 0002, Timothy M. Hospedales, Chen Change Loy |
CVPR | 1 |
| 2015 | Sketch-a-Net that Beats HumansabstractWe propose a multi-scale multi-channel deep neural network framework that, for the first time, yields sketch recognition performance surpassing that of humans. Our superior performance is a result of explicitly embedding the unique characteristics of sketches in our model: (i) a network architecture designed for sketch rather than natural photo statistics, (ii) a multi-channel generalisation that encodes sequential ordering in the sketching process, and (iii) a multi-scale network ensemble with joint Bayesian fusion that accounts for the different levels of abstraction exhibited in free-hand sketches. We show that state-of-the-art deep networks specifically engineered for photos of natural objects fail to perform well on sketch recognition, regardless whether they are trained using photo or sketch. Our network on the other hand not only delivers the best performance on the largest human sketch dataset to date, but also is small in size making efficient training possible using just CPUs. Qian Yu 0002, Yongxin Yang, Yi-Zhe Song, Tao Xiang 0002, Timothy M. Hospedales |
BMVC | 1 |