VLDB 2026 Research / reviewers in the wild / expert
Junyan Wang 0001
dblp:70/4949-1
· DBLP profile ↗
14ranked-venue papers
7as first author
12since 2021 · last 2026
0000-0001-5409-1292ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | B3CT: Three-Branch Learning with Unlabeled Target Signals for Domain-Robust Semantic Segmentation
Xin Zhao 0012, Jian Jia, Junyan Wang 0001, Lijun Cao |
Int. J. Comput. Vis. | 4 |
| 2025 | Distribution Optimization Under Gaussian Hypothesis for Domain Adaptive Semantic SegmentationabstractDomain adaptive semantic segmentation aims to transfer a model, proficient in dense image classification, from a source domain to a target domain. While various transfer methods have been explored in previous studies, we argue that the modeling of categories within the model significantly affects its transferability. Building on the Gaussian Hypothesis, which posits that each category in the feature space adheres to a multidimensional Gaussian distribution, we propose a Class-Aware Variational Inference (CAVI) training method. This approach normalizes features of different categories into distinct multidimensional Gaussian distributions. To further learn domain-independent feature distributions, we optimize the feature space using a Gaussian-based alignment strategy and incorporate Gaussian-based contrastive learning. Experimental results demonstrate that our method achieves state-of-the-art performance on the GTAV → Cityscapes and Synthia → Cityscapes benchmarks. Xin Zhao 0012, Junyan Wang 0001, Lijun Cao, Junge Zhang |
WACV | 4 |
| 2024 | Towards Effective Usage of Human-Centric Priors in Diffusion Models for Text-based Human Image GenerationabstractVanilla text-to-image diffusion models struggle with generating accurate human images, commonly resulting in imperfect anatomies such as unnatural postures or disproportionate limbs. Existing methods address this issue mostly by fine-tuning the model with extra images or adding additional control - human-centric priors such as pose or depth maps - during the image generation phase. This paper explores the integration of these human-centric priors directly into the model fine-tuning stage, essentially eliminating the need for extra conditions at the inference stage. We realize this idea by proposing a human-centric alignment loss to strengthen human-related information from the textual prompts within the cross-attention maps. To ensure semantic detail richness and human structural accuracy during fine-tuning, we introduce scale-aware and step-wise constraints within the diffusion process, according to an indepth analysis of the cross-attention layer. Extensive experiments show that our method largely improves over state-of-the-art text-to-image models to synthesize high-quality human images based on user-written prompts. Project page: https://hcplayercvpr2024.github.io. Junyan Wang 0001, Zhenhong Sun, Zhiyu Tan, Xuanbai Chen, Hao Li 0030, Cheng Zhang 0014, Yang Song 0001 |
CVPR | 1 |
| 2024 | EGGen: Image Generation with Multi-entity Prior Learning through Entity GuidanceabstractDiffusion models have shown remarkable prowess in text-to-image synthesis and editing, yet they often stumble when tasked with interpreting complex prompts that describe multiple entities with specific attributes and interrelations. The generated images often contain inconsistent multi-entity representation (IMR), reflected as inaccurate presentations of the multiple entities and their attributes. Although providing spatial layout guidance improves the multi-entity generation quality in existing works, it is still challenging to handle the leakage attributes and avoid unnatural characteristics. To address the IMR challenge, we first conduct in-depth analyses of the diffusion process and attention operation, revealing that the IMR challenges largely stem from the process of cross-attention mechanisms. According to the analyses, we introduce the entity guidance generation mechanism, which maintains the integrity of the original diffusion model parameters by integrating plug-in networks. Our work advances the stable diffusion model by segmenting comprehensive prompts into distinct entity-specific prompts with bounding boxes, enabling a transition from multi-entity to single-entity generation in cross-attention layers. More importantly, we introduce entity-centric cross-attention layers that focus on individual entities to preserve their uniqueness and accuracy, alongside global entity alignment layers that refine cross-attention maps using multi-entity priors for precise positioning and attribute accuracy. Additionally, a linear attenuation module is integrated to progressively reduce the influence of these layers during inference, preventing oversaturation and preserving generation fidelity. Our comprehensive experiments demonstrate that this entity guidance generation enhances existing text-to-image models in generating detailed, multi-entity images. Zhenhong Sun, Junyan Wang 0001, Zhiyu Tan, Daoyi Dong, Hailan Ma, Hao Li 0030, Dong Gong |
ACM Multimedia | 2 |
| 2024 | Deconfounding Causal Inference for Zero-Shot Action RecognitionabstractZero-shot action recognition (ZSAR) aims to recognize unseen action categories in the test set without corresponding training examples. Most existing zero-shot methods follow the feature generation framework to transfer knowledge from seen action categories to model the feature distribution of unseen categories. However, due to the complexity and diversity of actions, it remains challenging to generate unseen feature distribution, especially for the cross-dataset scenario when there is a potentially larger domain shift. This article proposes aDeconfoundingCaUSAlGAN (DeCalGAN) for generating unseen action video features with the following technical contributions: 1) Our model unifies compositional ZSAR with traditional visual-semantic models to incorporate local object information with global semantic information for feature generation. 2) A GAN-based architecture is proposed for causal inference and unseen distribution discovery. 3) A deconfounding module is proposed to refine representations of local objects and global semantic information confounder in the training data. Action descriptions and random object features after causal inference are then used to discover unseen distributions of novel actions in different datasets. Our extensive experiments onCross-DatasetZero-ShotActionRecognition (CD-ZSAR) demonstrate substantial improvement over the UCF101 and HMDB51 standard benchmarks for this problem. Junyan Wang 0001, Yiqi Jiang, Yang Long 0001, Xiuyu Sun, Maurice Pagnucco, Yang Song 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | Maximizing Spatio-Temporal Entropy of Deep 3D CNNs for Efficient Video Recognition
Junyan Wang 0001, Zhenhong Sun, Yichen Qian, Dong Gong, Xiuyu Sun, Ming Lin 0002, Maurice Pagnucco, Yang Song 0001 |
ICLR | 1 |
| 2023 | Community-Aware Federated Video SummarizationabstractVideo summarization aims to extract representative frames to retain high-level information. Increasing concerns about privacy issues have been raised because conventional large-scale training requires users to upload video samples that may inevitably release sensitive information. In this paper, we thoroughly discuss the Federated Video Summarization problem, i.e., how to obtain a robust video summarization model when video data is distributed on private data islands. Our key contribution includes 1) We propose a fundamental Frame-Based aggregation method to video-related tasks, which differs from the sample-based aggregation in conventional FedAvg. 2) To mitigate the heterogeneous distribution due to community diversity, we propose the Community-Aware Clustering Federated Video Summarization Framework (CFed-VS) that clusters clients via a novel data-driven clustering approach. 3) We further tackle the challenging non-IID setting with a proposed Mixture Transformer, which manifests state-of-the-art performance via extensive quantitative and qualitative experiments on TVSum and SumMe datasets. Fan Wan, Junyan Wang 0001, Haoran Duan 0001, Yang Song 0001, Maurice Pagnucco, Yang Long 0001 |
IJCNN | 2 |
| 2022 | Towards Unified Multi-Excitation for Unsupervised Video Prediction
Junyan Wang 0001, Likun Qin, Peng Zhang 0058, Yang Long 0001, Bingzhang Hu, Maurice Pagnucco, Shizheng Wang, Yang Song 0001 |
BMVC | 1 |
| 2022 | GiraffeDet: A Heavy-Neck Paradigm for Object Detection
Yiqi Jiang, Zhiyu Tan, Junyan Wang 0001, Xiuyu Sun, Ming Lin 0002, Hao Li 0030 |
ICLR | 3 |
| 2022 | Entropy-Driven Mixed-Precision Quantization for Deep Network DesignabstractDeploying deep convolutional neural networks on Internet-of-Things (IoT) devices is challenging due to the limited computational resources, such as limited SRAM memory and Flash storage. Previous works re-design a small network for IoT devices, and then compress the network size by mixed-precision quantization. This two-stage procedure cannot optimize the architecture and the corresponding quantization jointly, leading to sub-optimal tiny deep models. In this work, we propose a one-stage solution that optimizes both jointly and automatically. The key idea of our approach is to cast the joint architecture design and quantization as an Entropy Maximization process. Particularly, our algorithm automatically designs a tiny deep model such that: 1) Its representation capacity measured by entropy is maximized under the given computational budget; 2) Each layer is assigned with a proper quantization precision; 3) The overall design loop can be done on CPU, and no GPU is required. More impressively, our method can directly search high-expressiveness architecture for IoT devices within less than half a CPU hour. Extensive experiments on three widely adopted benchmarks, ImageNet, VWW and WIDER FACE, demonstrate that our method can achieve the state-of-the-art performance in the tiny deep model regime. Code and pre-trained models are available at https://github.com/alibaba/lightweight-neural-architecture-search. Zhenhong Sun, Ce Ge, Junyan Wang 0001, Ming Lin 0002, Hesen Chen, Hao Li 0030, Xiuyu Sun |
NeurIPS | 3 |
| 2021 | Dynamic Graph Warping Transformer for Video Alignment
Junyan Wang 0001, Yang Long 0001, Maurice Pagnucco, Yang Song 0001 |
BMVC | 1 |
| 2021 | Discriminative Latent Semantic Graph for Video CaptioningabstractVideo captioning aims to automatically generate natural language sentences that can describe the visual contents of a given video. Existing generative models like encoder-decoder frameworks cannot explicitly explore the object-level interactions and frame-level information from complex spatio-temporal data to generate semantic-rich captions. Our main contribution is to identify three key problems in a joint framework for future video summarization tasks. 1) Enhanced Object Proposal: we propose a novel Conditional Graph that can fuse spatio-temporal information into latent object proposal. 2) Visual Knowledge: Latent Proposal Aggregation is proposed to dynamically extract visual words with higher semantic levels. 3) Sentence Validation: A novel Discriminative Language Validator is proposed to verify generated captions so that key semantic concepts can be effectively preserved. Our experiments on two public datasets (MVSD and MSR-VTT) manifest significant improvements over state-of-the-art approaches on all metrics, especially for BLEU-4 and CIDEr. Our code is available at https://github.com/baiyang4/D-LSG-Video-Caption. Yang Bai 0011, Junyan Wang 0001, Yang Long 0001, Bingzhang Hu, Yang Song 0001, Maurice Pagnucco, Yu Guan 0001 |
ACM Multimedia | 2 |
| 2020 | Query Twice: Dual Mixture Attention Meta Learning for Video SummarizationabstractVideo summarization aims to select representative frames to retain high-level information, which is usually solved by predicting the segment-wise importance score via a softmax function. However, softmax function suffers in retaining high-rank representations for complex visual or sequential information, which is known as the Softmax Bottleneck problem. In this paper, we propose a novel framework named Dual Mixture Attention (DMASum) model with Meta Learning for video summarization that tackles the softmax bottleneck problem, where the Mixture of Attention layer (MoA) effectively increases the model capacity by employing twice self-query attention that can capture the second-order changes in addition to the initial query-key attention, and a novel Single Frame Meta Learning rule is then introduced to achieve more generalization to small datasets with limited training sources. Furthermore, the DMASum significantly exploits both visual and sequential attention that connects local key-frame and global attention in an accumulative way. We adopt the new evaluation protocol on two public datasets, SumMe, and TVSum. Both qualitative and quantitative experiments manifest significant improvements over the state-of-the-art methods. Junyan Wang 0001, Yang Bai 0011, Yang Long 0001, Bingzhang Hu, Zhenhua Chai, Yu Guan 0001, Xiaolin Wei |
ACM Multimedia | 1 |
| 2019 | Order Matters: Shuffling Sequence Generation for Video Prediction
Junyan Wang 0001, Bingzhang Hu, Yang Long 0001, Yu Guan 0001 |
BMVC | 1 |