EDBT 2026 Demo / reviewers in the wild / expert
Jing Zhang 0017
dblp:05/3499-17
· DBLP profile ↗
37ranked-venue papers
8as first author
30since 2021 · last 2026
0000-0003-3516-0111ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 26 · 4 first-author · 22 since 2021Artificial intelligence and machine learning · 16 · 4 first-author · 11 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Free Your Hands: Human-Demonstration-Free Diffusion Policy Adaptation for Robotic ManipulationabstractDiffusion-based policies demonstrate remarkable capabilities in generating robust and precise actions for embodied tasks. However, when adapting to a new environment with visual domain gap such as lighting variation, object appearance change, different background, and unseen distractors, these methods require a substantial amount of human demonstrations gathered through teleoperation. This data collection process incurs significant costs in terms of both time and financial resources. To this end, we introduceFreeHand, which enables cross-environment adaptation without requiring additional teleoperated demonstrations,freeing human handsfrom the time-consuming data collection process. Specifically, we find that a well-trained diffusion policy network struggles to adapt to new environments due to distribution mismatches at both the visual and policy levels. Simply aligning domains at the visual level is insufficient, as even subtle visual changes in the environment can lead to severe action failures. Therefore, we propose a two-level alignment scheme. The visual-level alignment is achieved through adversarial training between the visual features from old and new environments. For the policy-level alignment, we introduce a novel noise-aware adversarial learning strategy for the diffusion policy. We validate our approach through comprehensive cross-domain adaptation experiments under three settings: Sim2Sim (PushT, CALVIN, RoboTwin), Real2Sim (SimplerEnv), and Real2Real (real-world robotic deployment). Our method demonstrates significant improvement with minimal additional parameter overhead, showcasing its effectiveness and efficiency. Ge Yuan, Jing Zhang 0017, Jianxin Pang, Dong Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | ELBO-T2IAlign: A Generic ELBO-Based Method for Calibrating Pixel-Level Text-Image Alignment in Diffusion ModelsabstractDiffusion models excel at image generation. Recent studies have shown that these models not only generate high-quality images but also encode text-image alignment information through attention maps or loss functions. This information is valuable for various downstream tasks, including segmentation, text-guided image editing, and compositional image generation. However, current methods heavily rely on the assumption of perfect text-image alignment in diffusion models, which is not the case. In this paper, we propose using zero-shot referring image segmentation as a proxy task to evaluate the pixel-level image and class-level text alignment of popular diffusion models. We conduct an in-depth analysis of pixel-text misalignment in diffusion models from the perspective of training data bias. We find that misalignment occurs in images with small-sized, occluded, or rare object classes. Therefore, we propose ELBO-T2IAlign-a simple yet effective method to calibrate pixel-text alignment in diffusion models based on the evidence lower bound (ELBO) of likelihood. ELBO-T2IAlign is training-free and generic: it requires no additional annotations, model retraining, or architectural modifications, and it can be directly applied to different diffusion backbones. Extensive experiments on zero-shot referring image segmentation, text-guided image editing, and compositional image generation verify that the proposed calibration improves pixel-text alignment across complementary downstream tasks. Qin Zhou 0003, Jing Zhang 0017, Qian Yu 0002, Lu Sheng, Dong Xu 0001 |
IEEE Trans. Image Process. | 5 |
| 2025 | Empowering LLMs to Understand and Generate Complex Vector GraphicsabstractThe unprecedented advancements in Large Language Models (LLMs) have profoundly impacted natural language processing but have yet to fully embrace the realm of scalable vector graphics (SVG) generation. While LLMs encode partial knowledge of SVG data from web pages during training, recent findings suggest that semantically ambiguous and tokenized representations within LLMs may result in hallucinations in vector primitive predictions. Additionally, LLM training typically lacks modeling and understanding of the rendering sequence of vector paths, which can lead to occlusion between output vector primitives. In this paper, we present LLM4SVG, an initial yet substantial step toward bridging this gap by enabling LLMs to better understand and generate vector graphics. LLM4SVG facilitates a deeper understanding of SVG components through learnable semantic tokens, which precisely encode these tokens and their corresponding properties to generate semantically aligned SVG outputs. Using a series of learnable semantic tokens, a structured dataset for instruction following is developed to support comprehension and generation across two primary tasks. Our method introduces a modular architecture to existing large language models, integrating semantic tags, vector instruction encoders, fine-tuned commands, and powerful LLMs to tightly combine geometric, appearance, and language information. To overcome the scarcity of SVG-text instruction data, we developed an automated data generation pipeline that collected our SVGX-SFT Dataset, consisting of high-quality human-designed SVGs and 580k SVG instruction following data specifically crafted for LLM training, which facilitated the adoption of the supervised fine-tuning strategy popular in LLM development. By exploring various training strategies, we developed LLM4SVG, which significantly moves beyond optimized rendering-based approaches and language-model-based baselines to achieve remarkable results in human evaluation tasks. Code, model, and data will be released at: https://ximinng.github.io/LLM4SVGProject/ Ximing Xing, Juncheng Hu 0001, Guotao Liang, Jing Zhang 0017, Dong Xu 0001, Qian Yu 0002 |
CVPR | 4 |
| 2025 | VectorPainter: Advanced Stylized Vector Graphics Synthesis Using Stroke-Style PriorsabstractWe introduce VectorPainter, a novel framework designed for reference-guided text-to-vector-graphics synthesis. Based on our observation that the style of strokes can be an important aspect to distinguish different artists, our method reforms the task into synthesizing a desired vector graphic by rearranging stylized strokes, which are vectorized from the reference images. Specifically, our method first converts the pixels of the reference image into a series of vector strokes, and then generates a vector graphic based on the input text description by optimizing the positions and colors of these vector strokes. To precisely capture the style of the reference image in the vectorized strokes, we propose an innovative vectorization method that employs an imitation learning strategy. To preserve the style of the strokes throughout the generation process, we introduce a style-preserving loss function. Extensive experiments have been conducted to demonstrate the superiority of our approach over existing works in stylized vector graphics synthesis, as well as the effectiveness of the various components of our method. Code, model, and data will be released at: https://hjc-owo.github.io/VectorPainterProject/ Juncheng Hu 0001, Ximing Xing, Jing Zhang 0017, Qian Yu 0002 |
ICME | 3 |
| 2025 | Multi-Object Sketch Animation with Grouping and Motion Trajectory Priors
Guotao Liang, Juncheng Hu 0001, Ximing Xing, Jing Zhang 0017, Qian Yu 0002 |
ACM Multimedia | 4 |
| 2025 | CAD-Coder: Text-to-CAD Generation with Chain-of-Thought and Geometric RewardabstractIn this work, we introduce CAD-Coder, a novel framework that reformulates text-to-CAD as the generation of CadQuery scripts—a Python-based, parametric CAD language.
This representation enables direct geometric validation, a richer modeling vocabulary, and seamless integration with existing LLMs.
To further enhance code validity and geometric fidelity, we propose a two-stage learning pipeline: (1) supervised fine-tuning on paired text–CadQuery data, and (2) reinforcement learning with Group Reward Policy Optimization (GRPO), guided by a CAD-specific reward comprising both a geometric reward (Chamfer Distance) and a format reward.
We also introduce a chain-of-thought (CoT) planning process to improve model reasoning, and construct a large-scale, high-quality dataset of 110K text–CadQuery–3D model triplets and 1.5K CoT samples via an automated pipeline. Extensive experiments demonstrate that CAD-Coder enables LLMs to generate diverse, valid, and complex CAD models directly from natural language, advancing the state of the art of text-to-CAD generation and geometric reasoning. Yandong Guan, Xilin Wang, Ximing Xing, Jing Zhang 0017, Dong Xu 0001, Qian Yu 0002 |
NeurIPS | 4 |
| 2025 | SVGDreamer++: Advancing Editability and Diversity in Text-Guided SVG GenerationabstractRecently, text-guided scalable vector graphics (SVG) synthesis has shown great promise in domains like iconography and sketching. However, existing Text-to-SVG methods often face challenges in editability, visual quality, and diversity. To address these issues, we propose a novel framework for text-guided SVG synthesis that significantly enhances editability, quality, and diversity. To enhance the editability of output SVGs, we introduce a Hierarchical Image VEctorization (HIVE) framework that operates at the semantic object level and supervises the optimization of components within the vector object. This approach facilitates the decoupling of vector graphics into distinct objects and component levels. Our proposed HIVE algorithm, informed by image segmentation priors, not only ensures a more precise representation of vector graphics but also enables fine-grained editing capabilities within vector objects. To improve the diversity of output SVGs, we present a Vectorized Particle-based Score Distillation (VPSD) approach. VPSD addresses over-saturation issues in existing methods and enhances sample diversity. A pre-trained reward model is incorporated to re-weight vector particles, improving aesthetic appeal and enabling faster convergence. Additionally, we design a novel adaptive vector primitives control strategy, which allows for the dynamic adjustment of the number of primitives, thereby enhancing the presentation of graphic details. Extensive experiments validate the effectiveness of the proposed method, demonstrating its superiority over baseline methods in terms of editability, visual quality, and diversity. We also show that our new method supports up to six distinct vector styles, capable of generating high-quality vector assets suitable for stylized vector design and poster design. Ximing Xing, Qian Yu 0002, Chuang Wang 0008, Haitao Zhou, Jing Zhang 0017, Dong Xu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Diffusion Model is Secretly a Training-Free Open Vocabulary Semantic SegmenterabstractThe pre-trained text-image discriminative models, such as CLIP, has been explored for open-vocabulary semantic segmentation with unsatisfactory results due to the loss of crucial localization information and awareness of object shapes. Recently, there has been a growing interest in expanding the application of generative models from generation tasks to semantic segmentation. These approaches utilize generative models either for generating annotated data or extracting features to facilitate semantic segmentation. This typically involves generating a considerable amount of synthetic data or requiring additional mask annotations. To this end, we uncover the potential of generative text-to-image diffusion models (e.g., Stable Diffusion) as highly efficient open-vocabulary semantic segmenters, and introduce a novel training-free approach named DiffSegmenter. The insight is that to generate realistic objects that are semantically faithful to the input text, both the complete object shapes and the corresponding semantics are implicitly learned by diffusion models. We discover that the object shapes are characterized by the self-attention maps while the semantics are indicated through the cross-attention maps produced by the denoising U-Net, forming the basis of our segmentation results. Additionally, we carefully design effective textual prompts and a category filtering mechanism to further enhance the segmentation results. Extensive experiments on three benchmark datasets show that the proposed DiffSegmenter achieves impressive results for open-vocabulary semantic segmentation. Xiawei Li, Jing Zhang 0017, Qingyuan Xu, Qin Zhou 0003, Qian Yu 0002, Lu Sheng, Dong Xu 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | P2FTrack: Multi-Object Tracking with Motion Prior and Feature PosteriorabstractMultiple object tracking (MOT) has emerged as a crucial component of the rapidly developing computer vision. However, existing multi-object tracking methods often overlook the relationship between features and motion, hindering the ability to strike a performance balance between coupled motion and complex scenes. In this work, we propose a novel end-to-end multi-object tracking method that integrates motion and feature information. To achieve this, we introduce a motion prior generator that transforms motion information into attention masks. Additionally, we leverage prior-posterior fusion multi-head attention to combine the motion-derived priors and attention-based posteriors. Our proposed method is extensively evaluated on MOT17 and DanceTrack datasets through comprehensive experiments and ablation studies, demonstrating state-of-the-art performance in the feature-based method with reasonable speed. Hong Zhang 0018, Jiaxu Wan, Jing Zhang 0017, Ding Yuan 0001, Xuliang Li 0005, Yifan Yang 0003 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Multi-Modality Affinity Inference for Weakly Supervised 3D Semantic Segmentationabstract3D point cloud semantic segmentation has a wide range of applications. Recently, weakly supervised point cloud segmentation methods have been proposed, aiming to alleviate the expensive and laborious manual annotation process by leveraging scene-level labels. However, these methods have not effectively exploited the rich geometric information (such as shape and scale) and appearance information (such as color and texture) present in RGB-D scans. Furthermore, current approaches fail to fully leverage the point affinity that can be inferred from the feature extraction network, which is crucial for learning from weak scene-level labels. Additionally, previous work overlooks the detrimental effects of the long-tailed distribution of point cloud data in weakly supervised 3D semantic segmentation. To this end, this paper proposes a simple yet effective scene-level weakly supervised point cloud segmentation method with a newly introduced multi-modality point affinity inference module. The point affinity proposed in this paper is characterized by features from multiple modalities (e.g., point cloud and RGB), and is further refined by normalizing the classifier weights to alleviate the detrimental effects of long-tailed distribution without the need of the prior of category distribution. Extensive experiments on the ScanNet and S3DIS benchmarks verify the effectiveness of our proposed method, which outperforms the state-of-the-art by ~4% to ~ 6% mIoU. Codes are released at https://github.com/Sunny599/AAAI24-3DWSSG-MMA. Xiawei Li, Qingyuan Xu, Jing Zhang 0017, Qian Yu 0002, Lu Sheng, Dong Xu 0001 |
AAAI | 3 |
| 2024 | Data-Free Generalized Zero-Shot LearningabstractDeep learning models have the ability to extract rich knowledge from large-scale datasets. However, the sharing of data has become increasingly challenging due to concerns regarding data copyright and privacy. Consequently, this hampers the effective transfer of knowledge from existing data to novel downstream tasks and concepts. Zero-shot learning (ZSL) approaches aim to recognize new classes by transferring semantic knowledge learned from base classes. However, traditional generative ZSL methods often require access to real images from base classes and rely on manually annotated attributes, which presents challenges in terms of data restrictions and model scalability. To this end, this paper tackles a challenging and practical problem dubbed as data-free zero-shot learning (DFZSL), where only the CLIP-based base classes data pre-trained classifier is available for zero-shot classification. Specifically, we propose a generic framework for DFZSL, which consists of three main components. Firstly, to recover the virtual features of the base data, we model the CLIP features of base class images as samples from a von Mises-Fisher (vMF) distribution based on the pre-trained classifier. Secondly, we leverage the text features of CLIP as low-cost semantic information and propose a feature-language prompt tuning (FLPT) method to further align the virtual image features and textual features. Thirdly, we train a conditional generative model using the well-aligned virtual image features and corresponding semantic text features, enabling the generation of new classes features and achieve better zero-shot generalization. Our framework has been evaluated on five commonly used benchmarks for generalized ZSL, as well as 11 benchmarks for the base-to-new ZSL. The results demonstrate the superiority and effectiveness of our approach. Our code is available in https://github.com/ylong4/DFZSL. Jing Zhang 0017, Qian Yu 0002, Lu Sheng, Dong Xu 0001 |
AAAI | 2 |
| 2024 | SVGDreamer: Text Guided SVG Generation with Diffusion ModelabstractRecently, text-guided scalable vector graphics (SVGs) synthesis has shown promise in domains such as iconography and sketch. However, existing text-to-SVG generation methods lack editability and struggle with visual quality and result diversity. To address these limitations, we propose a novel text-guided vector graphics synthesis method called SVGDreamer. SVGDreamer incorporates a semantic-driven image vectorization (SIVE) process that enables the decomposition of synthesis into foreground objects and background, thereby enhancing editability. Specifically, the SIVE process introduces attention-based primitive control and an attention-mask loss function for effective control and manipulation of individual elements. Additionally, we propose a Vectorized Particle-based Score Distillation (VPSD) approach to address issues of shape over-smoothing, color over-saturation, limited diversity, and slow convergence of the existing text-to-SVG generation methods by modeling SVGs as distributions of control points and colors. Furthermore, VPSD leverages a reward model to re-weight vector particles, which improves aesthetic appeal and accelerates convergence. Extensive experiments are conducted to validate the effectiveness of SVGDreamer, demonstrating its superiority over baseline methods in terms of editability, visual quality, and diversity. Project page: https://ximinng.github.io/SVGDreamer-project/ Ximing Xing, Haitao Zhou, Chuang Wang 0008, Jing Zhang 0017, Dong Xu 0001, Qian Yu 0002 |
CVPR | 4 |
| 2024 | Dynamic Prompting of Frozen Text-to-Image Diffusion Models for Panoptic Narrative GroundingabstractPanoptic narrative grounding (PNG), whose core target is fine-grained image-text alignment, requires a panoptic segmentation of referred objects given a narrative caption. Previous discriminative methods achieve only weak or coarse-grained alignment by panoptic segmentation pretraining or CLIP model adaptation. Given the recent progress of text-to-image Diffusion models, several works have shown their capability to achieve fine-grained image-text alignment through cross-attention maps and improved general segmentation performance. However, the direct use of phrase features as static prompts to apply frozen Diffusion models to the PNG task still suffers from a large task gap and insufficient vision-language interaction, yielding inferior performance. Therefore, we propose an Extractive-Injective Phrase Adapter (EIPA) bypass within the Diffusion UNet to dynamically update phrase prompts with image features and inject the multimodal cues back, which leverages the fine-grained image-text alignment capability of Diffusion models more sufficiently. In addition, we also design a Multi-Level Mutual Aggregation (MLMA) module to reciprocally fuse multi-level image and phrase features for segmentation refinement. Extensive experiments on the PNG benchmark show that our method achieves new state-of-the-art performance. Tianrui Hui, Jing Zhang 0017, Bin Ma 0028, Xiaoming Wei, Jizhong Han, Si Liu 0001 |
ACM Multimedia | 4 |
| 2024 | 3D Reconstruction From a Single Sketch via View-Dependent Depth SamplingabstractReconstructing a 3D shape based on a single sketch image is challenging due to the inherent sparsity and ambiguity present in sketches. Existing methods lose fine details when extracting features to predict 3D objects from sketches. Upon analyzing the 3D-to-2D projection process, we observe that the density map, characterizing the distribution of 2D point clouds, can serve as a proxy to facilitate the reconstruction process. In this work, we propose a novel sketch-based 3D reconstruction model named SketchSampler. It initiates the process by translating a sketch through an image translation network into a more informative 2D representation, which is then used to generate a density map. Subsequently, a two-stage probabilistic sampling process is employed to reconstruct a 3D point cloud: first, recovering the 2D points (i.e., the x and y coordinates) by sampling the density map; and second, predicting the depth (i.e., the z coordinate) by sampling the depth values along the ray determined by each 2D point. Additionally, we convert the reconstructed point cloud into a 3D mesh for wider applications. To reduce ambiguity, we incorporate hidden lines in sketches. Experimental results demonstrate that our proposed approach significantly outperforms other baseline methods. Chenjian Gao, Xilin Wang, Qian Yu 0002, Lu Sheng, Jing Zhang 0017, Xiaoguang Han 0001, Yi-Zhe Song, Dong Xu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | SWGNet: Step-Wise Reference Frame Generation Network for Multiview Video CodingabstractIn multiview video coding, the coding performance highly depends on the quality of the reference frames. In view of this, a step-wise reference frame generation network (SWGNet) is designed to improve the quality of the reference frame for efficient multiview video coding. In particular, a frame-level to block-level learning paradigm is proposed to step-wisely generate a high-quality reference frame. In the frame-level stage, by exploiting parallax correlations between temporal and inter-view references on the basis of image alignment, a parallax-guided frame-level synthesis module is proposed to generate an elementary reference frame. Then, in the block-level stage, a transformer-based block-level aggregation module is designed to further refine the texture details of the reference frame by modeling long-range dependencies among pixels. The proposed SWGNet is integrated into 3D-HEVC, and extensive experiments demonstrate that the proposed method achieves significant bitrate saving compared with 3D-HEVC. Jing Zhang 0017, Yonghong Hou, Zhaoqing Pan, Bo Peng 0007, Nam Ling, Jianjun Lei 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Feature Decomposition for Reducing Negative Transfer: A Novel Multi-Task Learning Method for Recommender System (Student Abstract)abstractWe propose a novel multi-task learning method termed Feature Decomposition Network (FDN). The key idea of the proposed FDN is to reduce the phenomenon of feature redundancy by explicitly decomposing features into task-specific features and task-shared features with carefully designed constraints. Experimental results show that our proposed FDN can outperform the state-of-the-art (SOTA) methods by a noticeable margin on Ali-CCP. Jie Zhou 0029, Qian Yu 0002, Chuan Luo 0002, Jing Zhang 0017 |
AAAI | 4 |
| 2023 | PCHM-Net: A New Point Cloud Compression Framework for Both Human Vision and Machine VisionabstractRecently, point cloud data has attracted increasing attention in various machine vision tasks like classification and detection. However, directly transmitting the raw point cloud for such machine vision tasks will bring a huge bit-rate cost. In this work, we propose a new point cloud compression framework called PCHM-Net for both human vision and machine vision. Our proposed PCHM-Net adopts a two-branch structure with the shared octree-based compression module. To better compress the point cloud data and save bit-rate for machine vision tasks, we use the point cloud selection module to select a sparse set of points before octree construction, which allows us to use deeper octree structure and thus better reconstruct the point cloud coordinates for more discriminative feature extraction. We further propose a global feature aggregation-based classification module to deal with the sparse point cloud classification task. Comprehensive experiments on various point cloud benchmark datasets (e.g., ModelNet, ShapeNet and ScanNet) demonstrate that our newly proposed PCHM-Net achieves promising coding performance for both human vision and machine vision. Jing Zhang 0017 |
ICME | 3 |
| 2023 | Distortion-aware Transformer in 360° Salient Object DetectionabstractWith the emergence of VR and AR, 360° data attracts increasing attention from the computer vision and multimedia communities. Typically, 360° data is projected into 2D ERP (equirectangular projection) images for feature extraction. However, existing methods cannot handle the distortions that result from the projection, hindering the development of 360-data-based tasks. Therefore, in this paper, we propose a Transformer-based model called DATFormer to address the distortion problem. We tackle this issue from two perspectives. Firstly, we introduce two distortion-adaptive modules. The first is a Distortion Mapping Module, which guides the model to pre-adapt to distorted features globally. The second module is a Distortion-Adaptive Attention Block that reduces local distortions on multi-scale features. Secondly, to exploit the unique characteristics of 360° data, we present a learnable relation matrix and use it as part of the positional embedding to further improve performance. Extensive experiments are conducted on three public datasets, and the results show that our model outperforms existing 2D SOD (salient object detection) and 360 SOD methods. The source code is available at https://github.com/yjzhao19981027/DATFormer/. Yinjie Zhao, Lichen Zhao, Qian Yu 0002, Lu Sheng, Jing Zhang 0017, Dong Xu 0001 |
ACM Multimedia | 5 |
| 2023 | DiffSketcher: Text Guided Vector Sketch Synthesis through Latent Diffusion ModelsabstractEven though trained mainly on images, we discover that pretrained diffusion models show impressive power in guiding sketch synthesis. In this paper, we present DiffSketcher, an innovative algorithm that creates \textit{vectorized} free-hand sketches using natural language input. DiffSketcher is developed based on a pre-trained text-to-image diffusion model. It performs the task by directly optimizing a set of Bézier curves with an extended version of the score distillation sampling (SDS) loss, which allows us to use a raster-level diffusion model as a prior for optimizing a parametric vectorized sketch generator. Furthermore, we explore attention maps embedded in the diffusion model for effective stroke initialization to speed up the generation process. The generated sketches demonstrate multiple levels of abstraction while maintaining recognizability, underlying structure, and essential visual details of the subject drawn. Our experiments show that DiffSketcher achieves greater quality than prior work. The code and demo of DiffSketcher can be found at https://ximinng.github.io/DiffSketcher-project/. Ximing Xing, Chuang Wang 0008, Haitao Zhou, Jing Zhang 0017, Qian Yu 0002, Dong Xu 0001 |
NeurIPS | 4 |
| 2023 | Toward Explainable 3D Grounded Visual Question Answering: A New Benchmark and Strong BaselineabstractRecently, 3D vision-and-language tasks have attracted increasing research interest. Compared to other vision-and-language tasks, the 3D visual question answering (VQA) task is less exploited and is more susceptible to language priors and co-reference ambiguity. Meanwhile, a couple of recently proposed 3D VQA datasets do not well support 3D VQA task due to their limited scale and annotation methods. In this work, we formally define and address a 3D grounded question answering (GQA) task by collecting a new 3D VQA dataset, referred to as flexible and explainable 3D GQA (FE-3DGQA), with diverse and relatively free-form question-answer pairs, as well as dense and completely grounded bounding box annotations. To achieve more explainable answers, we label the objects appeared in the complex QA pairs with different semantic types, including answer-grounded objects (both appeared and not appeared in the questions), and contextual objects for answer-grounded objects. We also propose a new 3D VQA framework to effectively predict the completely visually grounded and explainable answer. Extensive experiments verify that our newly collected benchmark datasets can be effectively used to evaluate various 3D VQA methods from different aspects and our newly proposed framework also achieves the state-of-the-art performance on the new benchmark dataset. The datasets and the source code are available viahttps://github.com/zlccccc/3DVL_Codebase. Lichen Zhao, Daigang Cai, Jing Zhang 0017, Lu Sheng, Dong Xu 0001, Yinjie Zhao, Lipeng Wang 0005, Xibo Fan |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | 3DJCG: A Unified Framework for Joint Dense Captioning and Visual Grounding on 3D Point CloudsabstractObserving that the 3D captioning task and the 3D grounding task contain both shared and complementary information in nature, in this work, we propose a unified framework to jointly solve these two distinct but closely related tasks in a synergistic fashion, which consists of both shared task-agnostic modules and lightweight task-specific modules. On one hand, the shared task-agnostic modules aim to learn precise locations of objects, fine-grained attribute features to characterize different objects, and complex relations between objects, which benefit both captioning and visual grounding. On the other hand, by casting each of the two tasks as the proxy task of another one, the lightweight task-specific modules solve the captioning task and the grounding task respectively. Extensive experiments and ablation study on three 3D vision and language datasets demonstrate that our joint training frame-work achieves significant performance gains for each individual task and finally improves the state-of-the-art performance for both captioning and grounding tasks. Daigang Cai, Lichen Zhao, Jing Zhang 0017, Lu Sheng, Dong Xu 0001 |
CVPR | 3 |
| 2022 | Improving RGB-D Point Cloud Registration by Learning Multi-scale Local Linear Transformation
Xiaoliang Huo, Jing Zhang 0017, Lu Sheng, Dong Xu 0001 |
ECCV (32) | 4 |
| 2022 | Deep region segmentation-based intra prediction for depth video coding
Jing Zhang 0017, Yonghong Hou, Zhe Zhang 0041, Dengchao Jin, Peihan Zhang, Ge Li 0002 |
Multim. Tools Appl. | 1 |
| 2022 | VDM-DA: Virtual Domain Modeling for Source Data-Free Domain AdaptationabstractDomain adaptation aims to leverage a label-rich domain (the source domain) to help model learning in a label-scarce domain (the target domain). Most domain adaptation methods require the co-existence of source and target domain samples to reduce the distribution mismatch. However, access to the source domain samples may not always be feasible in real-world applications due to different problems (e.g., storage, transmission, and privacy issues). In this work, we deal with the source data-free unsupervised domain adaptation problem and propose a novel approach referred to as Virtual Domain Modeling for Domain Adaptation (VDM-DA), in which the virtual domain acts as a bridge between the source and target domains. Specifically, based on the pre-trained source model, we generate the virtual domain samples by using an approximated Gaussian Mixture Model (GMM) in the feature space, such that the virtual domain maintains a similar distribution with the source domain without access to the original source data. Moreover, we also design an effective distribution alignment method to reduce the distribution divergence between the virtual domain and the target domain by gradually improving the compactness of the target domain distribution through model learning. In this way, we successfully achieve the goal of distribution alignment between the source and target domains when training deep networks without access to the source domain data. We conduct extensive experiments on four benchmark datasets for both 2D image-based and 3D point cloud-based cross-domain object recognition tasks, where the proposed method referred to as Virtual Domain Modeling for Domain Adaptation (VDM-DA) achieves the promising performance on all datasets. Jing Zhang 0017, Wen Li 0001, Dong Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Cross-View Locality Preserved Diversity and Consensus Learning for Multi-View Unsupervised Feature SelectionabstractAlthough demonstrating great success, previous multi-view unsupervised feature selection (MV-UFS) methods often construct a view-specific similarity graph and characterize the local structure of data within each single view. In such a way, the cross-view information could be ignored. In addition, they usually assume that different feature views are projected from a latent feature space while the diversity of different views cannot be fully captured. In this work, we resent a MV-UFS model via cross-view local structure preserved diversity and consensus learning, referred to as CvLP-DCL briefly. In order to exploit both the shared and distinguishing information across different views, we project each view into a label space, which consists of a consensus part and a view-specific part. Therefore, we regularize the fact that different views represent same samples. Meanwhile, a cross-view similarity graph learning term with matrix-induced regularization is embedded to preserve the local structure of data in the label space. By imposing the$l_{2,1}$-norm on the feature projection matrices for constraining row sparsity, discriminative features can be selected from different views. An efficient algorithm is designed to solve the resultant optimization problem and extensive experiments on six publicly datasets are conducted to validate the effectiveness of the proposed CvLP-DCL. Chang Tang, Xinwang Liu 0002, Wei Zhang 0049, Jing Zhang 0017, Jian Xiong 0002, Lizhe Wang 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2021 | MMA-Net: A MultiModal-Attention-Based Deep Neural Network for Web Services Classification
Jing Zhang 0017, Changran Lei, Yilong Yang 0001, Borui Wang, Yang Chen 0062 |
ICSOC | 1 |
| 2021 | Transfer Learning for Web Services ClassificationabstractWeb service classification is one of the common approaches to discover and reuse services. Machine learning methods are widely used for web service classification. However, due to the limited high-quality services in the public dataset, the state-of-the-art deep learning methods can not achieve high accuracy. In this paper, we propose a transfer learning approach Tr-ServeNet to reuse the knowledge of the App classification problem for web service classification. We pre-train a deep learning model for the App classification problem, in which the dataset contains high-quality data from Apple Store, and then transfer the embedded and extracted features to assist web service classification. To demonstrate the effectiveness of our approach, we compare the proposed method with other existing machine learning methods on the 50-category benchmark with 10, 000 real-world web services. The experimental results indicate that the proposed transfer learning method can reach the highest Top-1 accuracy in the benchmark of service classification. Yilong Yang 0001, Zhaotian Li, Jing Zhang 0017, Yang Chen 0062 |
ICWS | 3 |
| 2021 | ServeNet-LT: A Normalized Multi-head Deep Neural Network for Long-tailed Web Services ClassificationabstractAutomatic service classification plays an important role in service discovery, selection, and composition. Recently, machine learning has been widely used in service classification. Though promising results are obtained, previous methods are merely evaluated on web services datasets with small-scale data and relatively balanced data, which limit their real-world applications. In this paper, we address the long-tailed web services classification problem with more categories and imbalanced data. Due to the long-tailed distribution of datasets, the existing machine learning and deep learning methods cannot work well. To deal with the long-tailed problem, we propose a normalized multi-head classifier learning strategy, which effectively reduces the classifier bias and benefit the generalization capacity of the extracted features. Extensive experiments are conducted on a large-scale long-tailed web services dataset, and the results show that our model outperforms the 11 compared service classification methods to a large margin. Jing Zhang 0017, Yang Chen 0062, Yilong Yang 0001, Changran Lei |
ICWS | 1 |
| 2021 | Source Data-free Unsupervised Domain Adaptation for Semantic SegmentationabstractDeep\footnote learning-based semantic segmentation methods require a huge amount of training images with pixel-level annotations. Unsupervised domain adaptation (UDA) for semantic segmentation enables transferring knowledge learned from the synthetic data (source domain) with low-cost annotations to the real images (target domain). However, current UDA methods mostly require full access to the source domain data for feasible adaptation, which limits their applications in real-world scenarios with privacy, storage, or transmission issues. To this end, this paper identifies and addresses a more practical but challenging problem of UDA for semantic segmentation, where access to the original source domain data is forbidden. In other words, only the pre-trained source model and unlabelled target domain data are available for adaptation. To tackle the problem, we propose to construct a set of source domain virtual data to mimic the source domain distribution by identifying the target domain high-confidence samples predicted by the pre-trained source model. Then by analyzing the data properties in the cross-domain semantic segmentation tasks, we propose an uncertainty and prior distribution-aware domain adaptation method to align the virtual source domain and the target domain with both adversarial learning and self-training strategies. Extensive experiments on three cross-domain semantic segmentation datasets with in-depth analyses verify the effectiveness of the proposed method. Mucong Ye, Jing Zhang 0017, Jinpeng Ouyang, Ding Yuan 0001 |
ACM Multimedia | 2 |
| 2021 | Progressive Modality Cooperation for Multi-Modality Domain AdaptationabstractIn this work, we propose a new generic multi-modality domain adaptation framework called Progressive Modality Cooperation (PMC) to transfer the knowledge learned from the source domain to the target domain by exploiting multiple modality clues (e.g., RGB and depth) under the multi-modality domain adaptation (MMDA) and the more general multi-modality domain adaptation using privileged information (MMDA-PI) settings. Under the MMDA setting, the samples in both domains have all the modalities. Through effective collaboration among multiple modalities, the two newly proposed modules in our PMC can select the reliable pseudo-labeled target samples, which captures the modality-specific information and modality-integrated information, respectively. Under the MMDA-PI setting, some modalities are missing in the target domain. Hence, to better exploit the multi-modality data in the source domain, we further propose the PMC with privileged information (PMC-PI) method by proposing a new multi-modality data generation (MMG) network. MMG generates the missing modalities in the target domain based on the source domain data by considering both domain distribution mismatch and semantics preservation, which are respectively achieved by using adversarial learning and conditioning on weighted pseudo semantic class labels. Extensive experiments on three image datasets and eight video datasets for various multi-modality cross-domain visual recognition tasks under both MMDA and MMDA-PI settings clearly demonstrate the effectiveness of our proposed PMC framework. Dong Xu 0001, Jing Zhang 0017, Wanli Ouyang |
IEEE Trans. Image Process. | 3 |
| 2020 | Enhanced Feature Pyramid Network for Semantic SegmentationabstractMulti-scale feature fusion has been an effective way for improving the performance of semantic segmentation. However, current methods generally fail to consider the semantic gaps between the shallow (low-level) and deep (high-level) features and thus the fusion methods may not be optimal. In this paper, to address the issues of the semantic gap between the feature from different layers, we propose a unified framework based on the U-shape encoder-decoder architecture, named Enhanced Feature Pyramid Network (EFPN). Specifically, the semantic enhancement module (SEM), edge extraction module (EEM), and context aggregation model (CAM) are incorporated into the decoder network to improve the robustness of the multilevel features aggregation. In addition, a global fusion model (GFM), which in the encoder branch is proposed to capture more semantic information in the deep layers and effectively transmit the high-level semantic features to each layer. Extensive experiments are conducted and the results show that the proposed framework achieves the state-of-the-art results on three public datasets, namely PASCAL VOC 2012, Cityscapes, and PASCAL Context. Furthermore, we also demonstrate that the proposed method is effective for other visual tasks that require frequent fusing features and upsampling. Mucong Ye, Jingpeng Ouyang, Jing Zhang 0017, Xiaogang Yu |
ICPR | 4 |
| 2019 | Unsupervised domain adaptation: A multi-task learning-based method
Jing Zhang 0017, Wanqing Li 0001, Philip Ogunbona |
Knowl. Based Syst. | 1 |
| 2018 | Importance Weighted Adversarial Nets for Partial Domain AdaptationabstractThis paper proposes an importance weighted adversarial nets-based method for unsupervised domain adaptation, specific for partial domain adaptation where the target domain has less number of classes compared to the source domain. Previous domain adaptation methods generally assume the identical label spaces, such that reducing the distribution divergence leads to feasible knowledge transfer. However, such an assumption is no longer valid in a more realistic scenario that requires adaptation from a larger and more diverse source domain to a smaller target domain with less number of classes. This paper extends the adversarial nets-based domain adaptation and proposes a novel adversarial nets-based partial domain adaptation method to identify the source samples that are potentially from the outlier classes and, at the same time, reduce the shift of shared classes between domains. Jing Zhang 0017, Zewei Ding, Wanqing Li 0001, Philip Ogunbona |
CVPR | 1 |
| 2017 | Joint Geometrical and Statistical Alignment for Visual Domain AdaptationabstractThis paper presents a novel unsupervised domain adaptation method for cross-domain visual recognition. We propose a unified framework that reduces the shift between domains both statistically and geometrically, referred to as Joint Geometrical and Statistical Alignment (JGSA). Specifically, we learn two coupled projections that project the source domain and target domain data into low-dimensional subspaces where the geometrical shift and distribution shift are reduced simultaneously. The objective function can be solved efficiently in a closed form. Extensive experiments have verified that the proposed method significantly outperforms several state-of-the-art domain adaptation methods on a synthetic dataset and three different real world cross-domain visual recognition tasks. Jing Zhang 0017, Wanqing Li 0001, Philip Ogunbona |
CVPR | 1 |
| 2016 | RGB-D-based action recognition datasets: A survey
Jing Zhang 0017, Wanqing Li 0001, Philip Ogunbona, Pichao Wang, Chang Tang |
Pattern Recognit. | 1 |
| 2016 | Action Recognition From Depth Maps Using Deep Convolutional Neural NetworksabstractThis paper proposes a new method, i.e., weighted hierarchical depth motion maps (WHDMM) + three-channel deep convolutional neural networks (3ConvNets), for human action recognition from depth maps on small training datasets. Three strategies are developed to leverage the capability of ConvNets in mining discriminative features for recognition. First, different viewpoints are mimicked by rotating the 3-D points of the captured depth maps. This not only synthesizes more data, but also makes the trained ConvNets view-tolerant. Second, WHDMMs at several temporal scales are constructed to encode the spatiotemporal motion patterns of actions into 2-D spatial structures. The 2-D spatial structures are further enhanced for recognition by converting the WHDMMs into pseudocolor images. Finally, the three ConvNets are initialized with the models obtained from ImageNet and fine-tuned independently on the color-coded WHDMMs constructed in three orthogonal planes. The proposed algorithm was evaluated on the MSRAction3D, MSRAction3DExt, UTKinect-Action, and MSRDailyActivity3D datasets using cross-subject protocols. In addition, the method was evaluated on the large dataset constructed from the above datasets. The proposed method achieved 2-9% better results on most of the individual datasets. Furthermore, the proposed method maintained its performance on the large dataset, whereas the performance of existing methods decreased with the increased number of actions. Pichao Wang, Wanqing Li 0001, Zhimin Gao, Jing Zhang 0017, Chang Tang, Philip Ogunbona |
IEEE Trans. Hum. Mach. Syst. | 4 |
| 2015 | ConvNets-Based Action Recognition from Depth Maps through Virtual Cameras and PseudocoloringabstractIn this paper, we propose to adopt ConvNets to recognize human actions from depth maps on relatively small datasets based on Depth Motion Maps (DMMs). In particular, three strategies are developed to effectively leverage the capability of ConvNets in mining discriminative features for recognition. Firstly, different viewpoints are mimicked by rotating virtual cameras around subject represented by the 3D points of the captured depth maps. This not only synthesizes more data from the captured ones, but also makes the trained ConvNets view-tolerant. Secondly, DMMs are constructed and further enhanced for recognition by encoding them into Pseudo-RGB images, turning the spatial-temporal motion patterns into textures and edges. Lastly, through transferring learning the models originally trained over ImageNet for image classification, the three ConvNets are trained independently on the color-coded DMMs constructed in three orthogonal planes. The proposed algorithm was extensively evaluated on MSRAction3D, MSRAction3DExt and UTKinect-Action datasets and achieved the state-of-the-art results on these datasets. Pichao Wang, Wanqing Li 0001, Zhimin Gao, Chang Tang, Jing Zhang 0017, Philip Ogunbona |
ACM Multimedia | 5 |