Jing Wang 0021

dblp:02/736-21 · DBLP profile ↗
← Back
18ranked-venue papers
11as first author
9since 2021 · last 2026
0000-0003-2065-1102ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 10 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2026 RelaCtrl: Relevance-Guided Efficient Control for Diffusion Transformers
abstract
The Diffusion Transformer plays a pivotal role in advancing text-to-image and text-to-video generation, owing primarily to its inherent scalability. However, existing controlled diffusion transformer methods incur significant parameter and computational overheads and suffer from inefficient resource allocation due to their failure to account for the varying relevance of control information across different transformer layers. To address this, we propose the Relevance-Guided Efficient Controllable Generation framework, RelaCtrl, enabling efficient and resource-optimized integration of control signals into the Diffusion Transformer. First, we evaluate the relevance of each layer in the Diffusion Transformer to the control information by assessing the ControlNet Relevance Score, which measures the impact of skipping each control layer on both the quality of generation and the control effectiveness during inference. Based on the strength of the relevance, we then tailor the positioning, parameter scale, and modeling capacity of the control layers to reduce unnecessary parameters and redundant computations. Additionally, to further improve efficiency, we replace the self-attention and FFN in the commonly used copy block with the carefully designed Two-Dimensional Shuffle Mixer (TDSM), enabling efficient implementation of both the token mixer and channel mixer. Both qualitative and quantitative experimental results demonstrate that our approach achieves superior performance with only 15% of the parameters and computational complexity compared to PixArt-delta.
Ke Cao 0001, Jing Wang 0021, Ao Ma 0005, Jiasong Feng, Xuanhua He, Run Ling, Haozhe Wang 0002, Hongjuan Pei, Yihua Shao, Zhanjie Zhang, Jie Zhang 0033
AAAI2
2025 PatchVSR: Breaking Video Diffusion Resolution Limits with Patch-wise Video Super-Resolution
abstract
Pre-trained video generation models hold great potential for generative video super-resolution (VSR). However, adapting them for full-size VSR, as most existing methods do, suffers from unnecessary intensive full-attention computation and fixed output resolution. To overcome these limitations, we make the first exploration into utilizing video diffusion priors for patch-wise VSR. This is non-trivial because pre-trained video diffusion models are not native for patch-level detail generation. To mitigate this challenge, we propose an innovative approach, called PatchVSR, which integrates a dual-stream adapter for conditional guidance. The patch branch extracts features from input patches to maintain content fidelity while the global branch extracts context features from the resized full video to bridge the generation gap caused by incomplete semantics of patches. Particularly, we also inject the patch’s location information into the model to better contextualize patch synthesis within the global video frame. Experiments demonstrate that our method can synthesize high-fidelity, high-resolution details at the patch level. A tailor-made multi-patch joint modulation is proposed to ensure visual consistency across individually enhanced patches. Due to the flexibility of our patch-based paradigm, we can achieve highly competitive 4K VSR based on a 512×512 resolution base model, with extremely high efficiency.
Shian Du, Menghan Xia, Chang Liu 0071, Xintao Wang 0002, Jing Wang 0021, Pengfei Wan 0001, Di Zhang 0026, Xiangyang Ji
CVPR5
2025 Lay2Story: Extending Diffusion Transformers for Layout-Togglable Story Generation
Ao Ma 0005, Jiasong Feng, Ke Cao 0001, Jing Wang 0021, Yun Wang 0053, Quanwei Zhang, Zhanjie Zhang
ICCV4
2025 PT-T2I/V: An Efficient Proxy-Tokenized Diffusion Transformer for Text-to-Image/Video-Task
abstract
The global self-attention mechanism in diffusion transformers involves redundant computation due to the sparse and redundant nature of visual information, and the attention map of tokens within a spatial window shows significant similarity. To address this redundancy, we propose the Proxy-Tokenized Diffusion Transformer (PT-DiT), which employs sparse representative token attention (where the number of representative tokens is much smaller than the total number of tokens) to efficiently model global visual information. Specifically, within each transformer block, we compute an averaging token from each spatial-temporal window to serve as a proxy token for that region. The global semantics are captured through the self-attention of these proxy tokens and then injected into all latent tokens via cross-attention. Simultaneously, we introduce window and shift window attention to address the limitations in detail modeling caused by the sparse attention mechanism. Building on the well-designed PT-DiT, we further develop the PT-T2I/V family, which includes a variety of models for T2I, T2V, and T2MV tasks. Experimental results show that PT-DiT achieves competitive performance while reducing computational complexity in image and video generation tasks (e.g., a reduction 59\% compared to DiT and a reduction 34\% compared to PixArt-$\alpha$). The visual exhibition of and code are available at https://360cvgroup.github.io/Qihoo-T2X/.
Jing Wang 0021, Ao Ma 0005, Jiasong Feng, Dawei Leng, Yuhui Yin, Xiaodan Liang
ICLR1
2025 FancyVideo: Towards Dynamic and Consistent Video Generation via Cross-frame Textual Guidance
abstract
Synthesizing motion-rich and temporally consistent videos remains a challenge in artificial intelligence, especially when dealing with extended durations. Existing text-to-video (T2V) models commonly employ spatial cross-attention for text control, equivalently guiding different frame generations without frame-specific textual guidance. Thus, the model's capacity to comprehend the temporal logic conveyed in prompts and generate videos with coherent motion is restricted. To tackle this limitation, we introduce FancyVideo, an innovative video generator that improves the existing text-control mechanism with the well-designed Cross-frame Textual Guidance Module (CTGM). Specifically, CTGM incorporates the Temporal Information Injector (TII) and Temporal Affinity Refiner (TAR) at the beginning and end of cross-attention, respectively, to achieve frame-specific textual guidance. Firstly, TII injects frame-specific information from latent features into text conditions, thereby obtaining cross-frame textual conditions. Then, TAR refines the correlation matrix between cross-frame textual conditions and latent features along the time dimension. Extensive experiments comprising both quantitative and qualitative evaluations demonstrate the effectiveness of FancyVideo. Our approach achieves state-of-the-art T2V generation results on the EvalCrafter benchmark and facilitates the synthesis of dynamic and consistent videos. Note that the T2V process of FancyVideo essentially involves a text-to-image step followed by T+I2V. This means it also supports the generation of videos from user images, i.e., the image-to-video (I2V) task. A significant number of experiments have shown that its performance is also outstanding.
Jiasong Feng, Ao Ma 0005, Jing Wang 0021, Ke Cao 0001, Zhanjie Zhang
IJCAI3
2025 WISA: World simulator assistant for physics-aware text-to-video generation
abstract
Recent advances in text-to-video (T2V) generation, exemplified by models such as Sora and Kling, have demonstrated strong potential for constructing world simulators. However, existing T2V models still struggle to understand abstract physical principles and to generate videos that faithfully obey physical laws. This limitation stems primarily from the lack of explicit physical guidance, caused by a significant gap between high-level physical concepts and the generative capabilities of current models. To address this challenge, we propose the **W**orld **S**imulator **A**ssistant (**WISA**), a novel framework designed to systematically decompose and integrate physical principles into T2V models. Specifically, WISA decomposes physical knowledge into three hierarchical levels: textual physical descriptions, qualitative physical categories, and quantitative physical properties. It then incorporates several carefully designed modules—such as Mixture-of-Physical-Experts Attention (MoPA) and a Physical Classifier—to effectively encode these attributes and enhance the model’s adherence to physical laws during generation. In addition, most existing video datasets feature only weak or implicit representations of physical phenomena, limiting their utility for learning explicit physical principles. To bridge this gap, we present **WISA-80K**, a new dataset comprising 80,000 human-curated videos that depict 17 fundamental physical laws across three core domains of physics: dynamics, thermodynamics, and optics. Experimental results show that WISA substantially improves the alignment of T2V models (such as CogVideoX and Wan2.1) with real-world physical laws, achieving notable gains on the VideoPhy benchmark. Our data, code, and models are available in the [Project Page](https://wisav1.github.io/WISA/).
Jing Wang 0021, Ao Ma 0005, Ke Cao 0001, Jiasong Feng, Zhanjie Zhang, Wanyuan Pang, Xiaodan Liang
NeurIPS1
2022 Metric-Fair Active Learning
abstract
Active learning has become a prevalent technique for designing label-efficient algorithms, where the central principle is to only query and fit “informative” labeled instances. It is, however, known that an active learning algorithm may incur unfairness due to such instance selection procedure. In this paper, we henceforth study metric-fair active learning of homogeneous halfspaces, and show that under the distribution-dependent PAC learning model, fairness and label efficiency can be achieved simultaneously. We further propose two extensions of our main results: 1) we show that it is possible to make the algorithm robust to the adversarial noise – one of the most challenging noise models in learning theory; and 2) it is possible to significantly improve the label complexity when the underlying halfspace is sparse.
Jie Shen 0005, Nan Cui, Jing Wang 0021
ICML3
2022 Partial Multi-Label Feature Selection
abstract
Multi-label feature selection is an effective approach to alleviate the high dimensionality of multi-label learning tasks. Most of the existing multi-label feature selection methods are based on the assumption that the relevant labels of each training sample are precisely annotated. However, in real-world applications, this assumption is not always held, each instance may be labeled with a set of candidate labels that contains all the ground-truth labels adulterated with noisy labels, which is called partial multi-label problem. Previous multi-label feature selection methods can not select the optimal feature set in the presence of partial multi-label. Therefore, to tackle this problem we propose a novel partial multi-label feature selection method, called PMLFS. To be specific, ground-truth labels and noisy labels are firstly distinguished in terms of label correlations and the sparsity of noisy labels. Secondly, manifold regularization is incorporated to explore the local structure. Based on the above, we design an objective framework involving$l_{2,1}$-norm regularization to achieve partial multi-label feature selection. Finally, extensive experiments on synthetic and real-world partial multi-label data sets demonstrate that the proposed method outperforms the state-of-the-art multi-label feature selection methods.
Jing Wang 0021, Pei-Pei Li 0001, Kui Yu
IJCNN1
2022 Fast spectral analysis for approximate nearest neighbor search
Jing Wang 0021, Jie Shen 0005
Mach. Learn.1
2018 Provable Variable Selection for Streaming Features
abstract
In large-scale machine learning applications and high-dimensional statistics, it is ubiquitous to address a considerable number of features among which many are redundant. As a remedy, online feature selection has attracted increasing attention in recent years. It sequentially reveals features and evaluates the importance of them. Though online feature selection has proven an elegant methodology, it is usually challenging to carry out a rigorous theoretical characterization. In this work, we propose a provable online feature selection algorithm that utilizes the online leverage score. The selected features are then fed to $k$-means clustering, making the clustering step memory and computationally efficient. We prove that with high probability, performing $k$-means clustering based on the selected feature space does not deviate far from the optimal clustering using the original data. The empirical results on real-world data sets demonstrate the effectiveness of our algorithm.
Jing Wang 0021, Jie Shen 0005, Ping Li 0001
ICML1
2018 A survey on online feature selection with streaming features
Xuegang Hu, Peng Zhou 0008, Pei-Pei Li 0001, Jing Wang 0021, Xindong Wu 0001
Frontiers Comput. Sci.4
2017 Online Matrix Completion for Signed Link Prediction
abstract
This work studies the binary matrix completion problem underlying a large body of real-world applications such as signed link prediction and information propagation. That is, each entry of the matrix indicates a binary preference such as "like" or "dislike", "trust" or "distrust". However, the performance of existing matrix completion methods may be hindered owing to three practical challenges: 1) the observed data are with binary label (i.e., not real value); 2) the data are typically sampled non-uniformly (i.e., positive links dominate the negative ones) and 3) a network may have a huge volume of data (i.e., memory and computational issue).
Jing Wang 0021, Jie Shen 0005, Ping Li 0001
WSDM1
2017 Object proposal with kernelized partial ranking
Jing Wang 0021, Jie Shen 0005, Ping Li 0001
Pattern Recognit.1
2015 Visual data denoising with a unified Schatten-p norm and ℓq norm regularized principal component pursuit
Jing Wang 0021, Meng Wang 0001, Xuegang Hu, Shuicheng Yan
Pattern Recognit.1
2015 Online Feature Selection with Group Structure Analysis
abstract
Online selection of dynamic features has attracted intensive interest in recent years. However, existing online feature selection methods evaluate features individually and ignore the underlying structure of a feature stream. For instance, in image analysis, features are generated in groups which represent color, texture, and other visual information. Simply breaking the group structure in feature selection may degrade performance. Motivated by this observation, we formulate the problem as an online group feature selection. The problem assumes that features are generated individually but there are group structures in the feature stream. To the best of our knowledge, this is the first time that the correlation among streaming features has been considered in the online feature selection process. To solve this problem, we develop a novel online group feature selection method named OGFS. Our proposed approach consists of two stages: online intra-group selection and online inter-group selection. In the intra-group selection, we design a criterion based on spectral analysis to select discriminative features in each group. In the inter-group selection, we utilize a linear regression model to select an optimal subset. This two-stage procedure continues until there are no more features arriving or some predefined stopping conditions are met. Finally, we apply our method to multiple tasks including image classification and face verification. Extensive empirical studies performed on real-world and benchmark data sets demonstrate that our method outperforms other state-of-the-art online feature selection methods.
Jing Wang 0021, Meng Wang 0001, Pei-Pei Li 0001, Luoqi Liu, Zhong-Qiu Zhao, Xuegang Hu, Xindong Wu 0001
IEEE Trans. Knowl. Data Eng.1
2014 Robust Face Recognition via Adaptive Sparse Representation
abstract
Sparse representation (or coding)-based classification (SRC) has gained great success in face recognition in recent years. However, SRC emphasizes the sparsity too much and overlooks the correlation information which has been demonstrated to be critical in real-world face recognition problems. Besides, some paper considers the correlation but overlooks the discriminative ability of sparsity. Different from these existing techniques, in this paper, we propose a framework called adaptive sparse representation-based classification (ASRC) in which sparsity and correlation are jointly considered. Specifically, when the samples are of low correlation, ASRC selects the most discriminative samples for representation, like SRC; when the training samples are highly correlated, ASRC selects most of the correlated and discriminative samples for representation, rather than choosing some related samples randomly. In general, the representation model is adaptive to the correlation structure that benefits from both l1-norm and l2-norm. Extensive experiments conducted on publicly available data sets verify the effectiveness and robustness of the proposed algorithm by comparing it with the state-of-the-art methods.
Jing Wang 0021, Canyi Lu, Meng Wang 0001, Pei-Pei Li 0001, Shuicheng Yan, Xuegang Hu
IEEE Trans. Cybern.1
2013 ApLeafis: An Android-Based Plant Leaf Identification System
Lin-Hai Ma, Zhong-Qiu Zhao, Jing Wang 0021
ICIC (1)3
2013 Online Group Feature Selection
Jing Wang 0021, Zhong-Qiu Zhao, Xuegang Hu, Yiu-Ming Cheung, Meng Wang 0001, Xindong Wu 0001
IJCAI1