Yuanzheng Ci

dblp:225/4774 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
4since 2021 · last 2023
0000-0001-5536-5355ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Representation and self-supervised learning · 37% Efficient and distributed learning · 20% Face, body and person analysis · 14%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%

Topics — the 14 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Face, body and person analysis
person re-identification
0.712023
HumanBench: Towards General Human-Centric Perception with Projector Assisted Pretraining · CVPR 2023
Machine learning › Representation and self-supervised learning
pre-training
0.712023
HumanBench: Towards General Human-Centric Perception with Projector Assisted Pretraining · CVPR 2023
Machine learning › Representation and self-supervised learning
contrastive learning
0.612022
Fast-MoCo: Boost Momentum-Based Contrastive Learning with Combinatorial Patches · ECCV (26) 2022
Machine learning › Representation and self-supervised learning › contrastive learning › self-supervised contrastive learning
momentum contrastive learning
0.612022
Fast-MoCo: Boost Momentum-Based Contrastive Learning with Combinatorial Patches · ECCV (26) 2022
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search
0.512021
Evolving Search Space for Neural Architecture Search · ICCV 2021
Machine learning › Efficient and distributed learning › automated machine learning › neural architecture search
search space design
0.512021
Evolving Search Space for Neural Architecture Search · ICCV 2021
Machine learning › Generative modeling › generative adversarial network
conditional GAN
0.312018
User-Guided Deep Anime Line Art Colorization with Conditional Adversarial Networks · ACM Multimedia 2018
Machine learning › Generative modeling
generative adversarial network
0.312018
User-Guided Deep Anime Line Art Colorization with Conditional Adversarial Networks · ACM Multimedia 2018
Visual content generation and editing
image colorization
0.312018
User-Guided Deep Anime Line Art Colorization with Conditional Adversarial Networks · ACM Multimedia 2018
Visual content generation and editing › image colorization
line art colorization
0.312018
User-Guided Deep Anime Line Art Colorization with Conditional Adversarial Networks · ACM Multimedia 2018
Computer vision › Segmentation and scene understanding
human parsing
0.212023
HumanBench: Towards General Human-Centric Perception with Projector Assisted Pretraining · CVPR 2023
Machine learning › Learning paradigms
multi-task learning
0.212023
UniHCP: A Unified Model for Human-Centric Perceptions · CVPR 2023
Computer vision › 3D vision
pose estimation
0.212023
HumanBench: Towards General Human-Centric Perception with Projector Assisted Pretraining · CVPR 2023
Machine learning › Optimization for machine learning
evolutionary computation
0.112021
Evolving Search Space for Neural Architecture Search · ICCV 2021

Methods — techniques the papers use, named apart from their topics

vision transformer · 0.7projector-assisted hierarchical pretraining · 0.7perceptual loss · 0.7large-scale joint training · 0.7WGAN-GP · 0.7momentum contrast · 0.6combinatorial patches · 0.6neural search-space evolution · 0.5multi-branch setting · 0.5
YearPublicationVenuePosition
2023 UniHCP: A Unified Model for Human-Centric Perceptions
abstract
Human-centric perceptions (e.g., pose estimation, human parsing, pedestrian detection, person re-identification, etc.) play a key role in industrial applications of visual models. While specific human-centric tasks have their own relevant semantic aspect to focus on, they also share the same underlying semantic structure of the human body. However, few works have attempted to exploit such homogeneity and design a general-propose model for human-centric tasks. In this work, we revisit a broad range of human-centric tasks and unify them in a minimalist manner. We propose UniHCP, a Unified Model for Human-Centric Perceptions, which unifies a wide range of human-centric tasks in a simplified end-to-end manner with the plain vision transformer architecture. With large-scale joint training on 33 human-centric datasets, UniHCP can outperform strong baselines on several in-domain and downstream tasks by direct evaluation. When adapted to a specific task, UniHCP achieves new SOTAs on a wide range of human-centric tasks, e.g., 69.8 mIoU on CIHP for human parsing, 86.18 mA on PA100K for attribute prediction, 90.3 mAP on Market1501 for ReID, and 85.8 JI on CrowdHuman for pedestrian detection, performing better than specialized models tailored for each task. The code and pretrained model are available at https://github.com/OpenGVLab/UniHCP.
Yuanzheng Ci, Yizhou Wang 0007, Meilin Chen, Shixiang Tang, Lei Bai 0001, Feng Zhu 0006, Rui Zhao 0001, Fengwei Yu, Donglian Qi, Wanli Ouyang
CVPR1
2023 HumanBench: Towards General Human-Centric Perception with Projector Assisted Pretraining
abstract
Human-centric perceptions include a variety of vision tasks, which have widespread industrial applications, including surveillance, autonomous driving, and the metaverse. It is desirable to have a general pretrain model for versatile human-centric downstream tasks. This paper forges ahead along this path from the aspects of both benchmark and pretraining methods. Specifically, we propose a HumanBench based on existing datasets to comprehensively evaluate on the common ground the generalization abilities of different pretraining methods on 19 datasets from 6 diverse downstream tasks, including person ReID, pose estimation, human parsing, pedestrian attribute recognition, pedestrian detection, and crowd counting. To learn both coarse-grained and fine-grained knowledge in human bodies, we further propose a Projector AssisTed Hierarchical pretraining method (PATH) to learn diverse knowledge at different granularity levels. Comprehensive evaluations on HumanBench show that our PATH achieves new state-of-the-art results on 17 downstream datasets and on-par results on the other 2 datasets. The code will be publicly at https://github.com/OpenGVLab/HumanBench.
Shixiang Tang, Qingsong Xie, Meilin Chen, Yizhou Wang 0007, Yuanzheng Ci, Lei Bai 0001, Feng Zhu 0006, Haiyang Yang, Rui Zhao 0001, Wanli Ouyang
CVPR6
2022 Fast-MoCo: Boost Momentum-Based Contrastive Learning with Combinatorial Patches
Yuanzheng Ci, Chen Lin 0003, Lei Bai 0001, Wanli Ouyang
ECCV (26)1
2021 Evolving Search Space for Neural Architecture Search
abstract
Automation of neural architecture design has been a coveted alternative to human experts. Various search methods have been proposed aiming to find the optimal architecture in the search space. One would expect the search results to improve when the search space grows larger since it would potentially contain more performant candidates. Surprisingly, we observe that enlarging search space is unbeneficial or even detrimental to existing NAS methods such as DARTS, ProxylessNAS, and SPOS. This counterintuitive phenomenon suggests that enabling existing methods to large search space regimes is non-trivial. However, this problem is less discussed in the literature.We present a Neural Search-space Evolution (NSE) scheme, the first neural architecture search scheme designed especially for large space neural architecture search problems. The necessity of a well-designed search space with constrained size is a tacit consent in existing methods, and our NSE aims at minimizing such necessity. Specifically, the NSE starts with a search space subset, then evolves the search space by repeating two steps: 1) search an optimized space from the search space subset, 2) refill this subset from a large pool of operations that are not traversed. We further extend the flexibility of obtainable architectures by introducing a learnable multi-branch setting. With the proposed method, we achieve 77.3% top-1 retrain accuracy on ImageNet with 333M FLOPs, which yielded a state-of-the-art performance among previous auto-generated architectures that do not involve knowledge distillation or weight pruning. When the latency constraint is adopted, our result also performs better than the previous best-performing mobile models with a 77.9% Top-1 retrain accuracy. Code is available at https://github.com/orashi/NSENAS.
Yuanzheng Ci, Chen Lin 0003, Ming Sun 0008, Hongwen Zhang 0001, Wanli Ouyang
ICCV1
2018 User-Guided Deep Anime Line Art Colorization with Conditional Adversarial Networks
abstract
Scribble colors based line art colorization is a challenging computer vision problem since neither greyscale values nor semantic information is presented in line arts, and the lack of authentic illustration-line art training pairs also increases difficulty of model generalization. Recently, several Generative Adversarial Nets (GANs) based methods have achieved great success. They can generate colorized illustrations conditioned on given line art and color hints. However, these methods fail to capture the authentic illustration distributions and are hence perceptually unsatisfying in the sense that they often lack accurate shading. To address these challenges, we propose a novel deep conditional adversarial architecture for scribble based anime line art colorization. Specifically, we integrate the conditional framework with WGAN-GP criteria as well as the perceptual loss to enable us to robustly train a deep network that makes the synthesized images more natural and real. We also introduce a local features network that is independent of synthetic data. With GANs conditioned on features from such network, we notably increase the generalization capability over "in the wild" line arts. Furthermore, we collect two datasets that provide high-quality colorful illustrations and authentic line arts for training and benchmarking. With the proposed model trained on our illustration dataset, we demonstrate that images synthesized by the presented approach are considerably more realistic and precise than alternative approaches.
Yuanzheng Ci, Xinzhu Ma, Zhihui Wang 0001, Zhongxuan Luo
ACM Multimedia1