EDBT 2026 Demo / reviewers in the wild / expert
Zhicheng Yan 0001
dblp:25/7295-1
· DBLP profile ↗
30ranked-venue papers
5as first author
14since 2021 · last 2025
0009-0005-0214-7916ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 4 first-author · 11 since 2021Artificial intelligence and machine learning · 20 · 3 first-author · 12 since 2021Databases, data management, data science and information retrieval · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MV-DUSt3R+: Single-Stage Scene Reconstruction from Sparse Views In 2 SecondsabstractRecent sparse multi-view scene reconstruction advances like DUSt3R and MASt3R no longer require camera calibration and camera pose estimation. However, they only process a pair of views at a time to infer pixel-aligned pointmaps. When dealing with more than two views, a combinatorial number of error prone pairwise reconstructions are usually followed by an expensive global optimization, which often fails to rectify the pairwise reconstruction errors. To handle more views, reduce errors, and improve inference time, we propose the fast single-stage feed-forward network MV- DUSt3R. At its core are multi-view decoder blocks which exchange information across any number of views while considering one reference view. To make our method robust to reference view selection, we further propose MV-DUSt3R+, which employs cross-reference-view blocks to fuse information across different reference view choices. To further enable novel view synthesis, we extend both by adding and jointly training Gaussian splatting heads. Experiments on multi-view stereo reconstruction, multi-view pose estimation, and novel view synthesis confirm that our methods improve significantly upon prior art. Code released.1 Zhenggang Tang, Yuchen Fan 0001, Dilin Wang, Hongyu Xu, Alexander G. Schwing, Zhicheng Yan 0001 |
CVPR | 7 |
| 2025 | DynamicVerse: A Physically-Aware Multimodal Framework for 4D World ModelingabstractUnderstanding the dynamic physical world, characterized by its evolving 3D structure, real-world motion, and semantic content with textual descriptions, is crucial for human-agent interaction and enables embodied agents to perceive and act within real environments with human‑like capabilities. However, existing datasets are often derived from limited simulators or utilize traditional Structure-from-Motion for up-to-scale annotation and offer limited descriptive captioning, which restricts the capacity of foundation models to accurately interpret real-world dynamics from monocular videos, commonly sourced from the internet. To bridge these gaps, we introduce **DynamicVerse**, a physical‑scale, multimodal 4D world modeling framework for dynamic real-world video. We employ large vision, geometric, and multimodal models to interpret metric-scale static geometry, real-world dynamic motion, instance-level masks, and holistic descriptive captions. By integrating window-based Bundle Adjustment with global optimization, our method converts long real-world video sequences into a comprehensive 4D multimodal format. DynamicVerse delivers a large-scale dataset consists of 100K+ videos with 800K+ annotated masks and 10M+ frames from internet videos. Experimental evaluations on three benchmark tasks, namely video depth estimation, camera pose estimation, and camera intrinsics estimation, demonstrate that our 4D modeling achieves superior performance in capturing physical-scale measurements with greater global accuracy than existing methods. Kairun Wen, Yuzhi Huang, Runyu Chen, Hui Zheng 0003, Yunlong Lin, Panwang Pan, Chenxin Li, Wenyan Cong, Junbin Lu, Chenguo Lin, Dilin Wang, Zhicheng Yan 0001, Hongyu Xu, Justin Theiss, Yue Huang 0001, Xinghao Ding, Zhiwen Fan |
NeurIPS | 13 |
| 2024 | Understanding and Accelerating Neural Architecture Search With Training-Free and Theory-Grounded MetricsabstractThis work targets designing a principled and unified training-free framework for Neural Architecture Search (NAS), with high performance, low cost, and in-depth interpretation. NAS has been explosively studied to automate the discovery of top-performer neural networks, but suffers from heavy resource consumption and often incurs search bias due to truncated training or approximations. Recent NAS works Mellor et al. 2021, Chen et al. 2021, Abdelfattah et al. 2021 start to explore indicators that can predict a network's performance without training. However, they either leveraged limited properties of deep networks, or the benefits of their training-free indicators were not applied to more extensive search methods. By rigorous correlation analysis, we present a unified framework to understand and accelerate NAS, by disentangling “TEG” characteristics of searched networks –Trainability,Expressivity,Generalization– all assessed in a training-free manner. The TEG indicators could be scaled up and integrated with various NAS search methods, including both supernet and single-path NAS approaches. Extensive studies validate the effective and efficient guidance from our TEG-NAS framework, leading to both improved search accuracy and over 56% reduction in search time cost. Moreover, we visualize search trajectories on three landscapes of “TEG” characteristics, observing that a good local minimum is easier to find on NAS-Bench-201 given its simple topology, whereas balancing “TEG” characteristics is much harder on the DARTS space due to its complex landscape geometry. Wuyang Chen 0001, Xinyu Gong, Yunchao Wei, Humphrey Shi, Zhicheng Yan 0001, Yi Yang 0001, Zhangyang Wang |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Exploring Open-Vocabulary Semantic Segmentation from CLIP Vision Encoder Distillation OnlyabstractSemantic segmentation is a crucial task in computer vision that involves segmenting images into semantically meaningful regions at the pixel level. However, existing approaches often rely on expensive human annotations as supervision for model training, limiting their scalability to large, unlabeled datasets. To address this challenge, we present ZeroSeg, a novel method that leverages the existing pretrained vision-language (VL) model (e.g. CLIP vision encoder [39]) to train open-vocabulary zero-shot semantic segmentation models. Although acquired extensive knowledge of visual concepts, it is non-trivial to exploit knowledge from these VL models to the task of semantic segmentation, as they are usually trained at an image level. ZeroSeg overcomes this by distilling the visual concepts learned by VL models into a set of segment tokens, each summarizing a localized region of the target image. We evaluate ZeroSeg on multiple popular segmentation benchmarks, including PASCAL VOC 2012, PASCAL Context, and COCO, in a zero-shot manner Our approach achieves state-of-the-art performance when compared to other zero-shot segmentation methods under the same training data, while also performing competitively compared to strongly supervised methods. Finally, we also demonstrated the effectiveness of ZeroSeg on open-vocabulary segmentation, through both human studies and qualitative visualizations. The code is publicly available at https://github.com/facebookresearch/ZeroSeg Jun Chen 0021, Deyao Zhu, Guocheng Qian, Bernard Ghanem, Zhicheng Yan 0001, Chenchen Zhu, Fanyi Xiao, Sean Culatana, Mohamed Elhoseiny 0001 |
ICCV | 5 |
| 2023 | Going Denser with Open-Vocabulary Part SegmentationabstractObject detection has been expanded from a limited number of categories to open vocabulary. Moving forward, a complete intelligent vision system requires understanding more fine-grained object descriptions, object parts. In this paper, we propose a detector with the ability to predict both open-vocabulary objects and their part segmentation. This ability comes from two designs. First, we train the detector on the joint of part-level, object-level and image-level data to build the multi-granularity alignment between language and image. Second, we parse the novel object into its parts by its dense semantic correspondence with the base object. These two designs enable the detector to largely benefit from various data sources and foundation models. In open-vocabulary part segmentation experiments, our method outperforms the baseline by 3.3∼7.3 mAP in cross-dataset generalization on PartImageNet, and improves the baseline by 7.3 novel AP50in cross-category generalization on Pascal Part. Finally, we train a detector that generalizes to a wide range of part segmentation datasets while achieving better performance than dataset-specific training. Peize Sun, Shoufa Chen, Chenchen Zhu, Fanyi Xiao, Ping Luo 0002, Saining Xie, Zhicheng Yan 0001 |
ICCV | 7 |
| 2023 | EgoObjects: A Large-Scale Egocentric Dataset for Fine-Grained Object UnderstandingabstractObject understanding in egocentric visual data is arguably a fundamental research topic in egocentric vision. However, existing object datasets are either non-egocentric or have limitations in object categories, visual content, and annotation granularities. In this work, we introduce EgoObjects, a large-scale egocentric dataset for fine-grained object understanding. Its Pilot version contains over 9K videos collected by 250 participants from 50+ countries using 4 wearable devices, and over 650K object annotations from 368 object categories. Unlike prior datasets containing only object category labels, EgoObjects also annotates each object with an instance-level identifier, and includes over 14K unique object instances. EgoObjects was designed to capture the same object under diverse background complexities, surrounding objects, distance, lighting and camera motion. In parallel to the data collection, we conducted data annotation by developing a multistage federated annotation process to accommodate the growing nature of the dataset. To bootstrap the research on EgoObjects, we present a suite of 4 benchmark tasks around the egocentric object understanding, including a novel instance level- and the classical category level object detection. Moreover, we also introduce 2 novel continual learning object detection tasks. The dataset and API are available at https://github.com/facebookresearch/EgoObjects. Chenchen Zhu, Fanyi Xiao, Andres Alvarado, Yasmine Babaei, Jiabo Hu, Hichem El-Mohri, Sean Culatana, Roshan Sumbaly, Zhicheng Yan 0001 |
ICCV | 9 |
| 2022 | Unified Transformer Tracker for Object TrackingabstractAs an important area in computer vision, object tracking has formed two separate communities that respectively study Single Object Tracking (SOT) and Multiple Object Tracking (MOT). However, current methods in one tracking scenario are not easily adapted to the other due to the divergent training datasets and tracking objects of both tasks. Although UniTrack [45] demonstrates that a shared appearance model with multiple heads can be used to tackle individual tracking tasks, it fails to exploit the large-scale tracking datasets for training and performs poorly on the single object tracking. In this work, we present the Unified Transformer Tracker (UTT) to address tracking problems in different scenarios with one paradigm. A track transformer is developed in our UTT to track the target in both SOT and MOT where the correlation between the target feature and the tracking frame feature is exploited to localize the target. We demonstrate that both SOT and MOT tasks can be solved within this framework, and the model can be simultaneously end-to-end trained by alternatively optimizing the SOT and MOT objectives on the datasets of individual tasks. Extensive experiments are conducted on several benchmarks with a unified model trained on both SOT and MOT datasets. Fan Ma, Zheng Shou 0001, Linchao Zhu, Haoqi Fan 0001, Yilei Xu, Yi Yang 0001, Zhicheng Yan 0001 |
CVPR | 7 |
| 2022 | NASViT: Neural Architecture Search for Efficient Vision Transformers with Gradient Conflict aware Supernet Training
Chengyue Gong, Dilin Wang, Meng Li 0004, Xinlei Chen, Zhicheng Yan 0001, Yuandong Tian, Qiang Liu 0001, Vikas Chandra |
ICLR | 5 |
| 2022 | Auto-X3D: Ultra-Efficient Video Understanding via Finer-Grained Neural Architecture SearchabstractEfficient video architecture is the key to deploying video recognition systems on devices with limited computing resources. Unfortunately, existing video architectures are often computationally intensive and not suitable for such applications. The recent X3D work presents a new family of efficient video models by expanding a hand-crafted image architecture along multiple axes, such as space, time, width, and depth. Although operating in a conceptually large space, X3D searches one axis at a time, and merely explored a small set of 30 architectures in total, which does not sufficiently explore the space. This paper bypasses existing 2D architectures, and directly searched for 3D architectures in a fine-grained space, where block type, filter number, expansion ratio and attention block are jointly searched. A probabilistic neural architecture search method is adopted to efficiently search in such a large space. Evaluations on Kinetics and Something-Something-V2 benchmarks confirm our AutoX3D models outperform existing ones in accuracy up to 1.3% under similar FLOPs, and reduce the computational cost up to ×1.74 when reaching similar performance. Yifan Jiang 0001, Xinyu Gong, Humphrey Shi, Zhicheng Yan 0001, Zhangyang Wang |
WACV | 5 |
| 2021 | FP-NAS: Fast Probabilistic Neural Architecture SearchabstractDifferential Neural Architecture Search (NAS) requires all layer choices to be held in memory simultaneously; this limits the size of both search space and final architecture. In contrast, Probabilistic NAS, such as PARSEC, learns a distribution over high-performing architectures, and uses only as much memory as needed to train a single model. Nevertheless, it needs to sample many architectures, making it computationally expensive for searching in an extensive space. To solve these problems, we propose a sampling method adaptive to the distribution entropy, drawing more samples to encourage explorations at the beginning, and reducing samples as learning proceeds. Furthermore, to search fast in the multivariate space, we propose a coarse-to-fine strategy by using a factorized distribution at the beginning which can reduce the number of architecture parameters by over an order of magnitude. We call this method Fast Probabilistic NAS (FP-NAS). Compared with PARSEC, it can sample 64% fewer architectures and search 2.1× faster. Compared with FBNetV2, FP-NAS is 1.9× - 3.5 ×faster, and the searched models outperform FBNetV2 models on ImageNet. FP-NAS allows us to expand the giant FBNetV2 space to be wider (i.e. larger channel choices) and deeper (i.e. more blocks), while adding Split-Attention block and enabling the search over the number of splits. When searching a model of size 0.4G FLOPS, FP-NAS is 132× faster than EfficientNet, and the searched FP-NAS-L0 model outperforms EfficientNet-B0 by 0.7% accuracy. Without using any architecture surrogate or scaling tricks, we directly search large models up to 1.0G FLOPS. Our FP-NAS-L2 model with simple distillation out-performs BigNAS-XL with advanced inplace distillation by 0.7% accuracy using similar FLOPS. Zhicheng Yan 0001, Xiaoliang Dai, Peizhao Zhang, Yuandong Tian, Bichen Wu, Matt Feiszli |
CVPR | 1 |
| 2021 | Multiscale Vision TransformersabstractWe present Multiscale Vision Transformers (MViT) for video and image recognition, by connecting the seminal idea of multiscale feature hierarchies with transformer models. Multiscale Transformers have several channel-resolution scale stages. Starting from the input resolution and a small channel dimension, the stages hierarchically expand the channel capacity while reducing the spatial resolution. This creates a multiscale pyramid of features with early layers operating at high spatial resolution to model simple low-level visual information, and deeper layers at spatially coarse, but complex, high-dimensional features. We evaluate this fundamental architectural prior for modeling the dense nature of visual signals for a variety of video recognition tasks where it outperforms concurrent vision transformers that rely on large scale external pre-training and are 5-10× more costly in computation and parameters. We further remove the temporal dimension and apply our model for image classification where it outperforms prior work on vision transformers. Code is available at: https://github.com/facebookresearch/SlowFast. Haoqi Fan 0001, Karttikeya Mangalam, Yanghao Li, Zhicheng Yan 0001, Jitendra Malik, Christoph Feichtenhofer |
ICCV | 5 |
| 2021 | Searching for Two-Stream Models in Multivariate Space for Video RecognitionabstractConventional video models rely on a single stream to capture the complex spatial-temporal features. Recent work on two-stream video models, such as SlowFast network and AssembleNet, prescribe separate streams to learn complementary features, and achieve stronger performance. However, manually designing both streams as well as the in-between fusion blocks is a daunting task, requiring to explore a tremendously large design space. Such manual exploration is time-consuming and often ends up with suboptimal architectures when computational resources are limited and the exploration is insufficient. In this work, we present a pragmatic neural architecture search approach, which is able to search for two-stream video models in giant spaces efficiently. We design a multivariate search space, including 6 search variables to capture a wide variety of choices in designing two-stream models. Furthermore, we propose a progressive search procedure, by searching for the architecture of individual streams, fusion blocks and attention blocks one after the other. We demonstrate two-stream models with significantly better performance can be automatically discovered in our design space. Our searched two-stream models, namely Auto-TSNet, consistently outperform other models on standard benchmarks. On Kinetics, compared with the SlowFast model, our Auto-TSNet-L model reduces FLOPS by nearly 11× while achieving the same accuracy 78.9%. On Something-Something-V2, Auto- TSNet-M improves the accuracy by at least 2% over other methods which use less than 50 GFLOPS per video. Xinyu Gong, Zheng Shou 0001, Matt Feiszli, Zhangyang Wang, Zhicheng Yan 0001 |
ICCV | 6 |
| 2021 | Visual Transformers: Where Do Transformers Really Belong in Vision Models?abstractA recent trend in computer vision is to replace convolutions with transformers. However, the performance gain of transformers is attained at a steep cost, requiring GPU years and hundreds of millions of samples for training. This excessive resource usage compensates for a misuse of transformers: Transformers densely model relationships between its inputs - ideal for late stages of a neural network, when concepts are sparse and spatially-distant, but extremely inefficient for early stages of a network, when patterns are redundant and localized. To address these issues, we leverage the respective strengths of both operations, building convolution-transformer hybrids. Critically, in sharp contrast to pixel-space transformers, our Visual Transformer (VT) operates in a semantic token space, judiciously attending to different image parts based on context. Our VTs significantly outperforms baselines: On ImageNet, our VT-ResNets outperform convolution-only ResNet by 4.6 to 7 points and transformer-only ViT-B by 2.6 points with 2.5× fewer FLOPs, 2.1× fewer parameters. For semantic segmentation on LIP and COCO-stuff, VT-based feature pyramid networks (FPN) achieve 0.35 points higher mIoU while reducing the FPN module’s FLOPs by 6.5x. Bichen Wu, Chenfeng Xu, Xiaoliang Dai, Alvin Wan, Peizhao Zhang, Zhicheng Yan 0001, Masayoshi Tomizuka, Joseph Gonzalez 0001, Kurt Keutzer, Peter Vajda |
ICCV | 6 |
| 2021 | Only Time Can Tell: Discovering Temporal Data for Temporal ModelingabstractUnderstanding temporal information and how the visual world changes over time, is a fundamental ability of intelligent systems. In video understanding, temporal information is at the core of many current challenges, including compression, efficient inference, motion estimation or summarization. However, in current video datasets it has been observed that action classes can often be recognized without any temporal information, from a single frame of video. As a result, both benchmarking and training in these datasets may give an unintentional advantage to models with strong image understanding capabilities, as opposed to those with strong temporal understanding. In other words, current datasets may not reward good temporal understanding, potentially hindering progress. In this paper we address this problem head on by identifying action classes where temporal information is actually necessary to recognize them and call these "temporal classes". Selecting temporal classes using a computational method would bias the process. Instead, we propose a methodology based on a simple and effective human annotation experiment. We remove just the temporal information, by shuffling frames in time, and measure if the action can still be recognized. Classes that cannot be recognized when frames are not in order, are included in the temporal set. We observe that this set is statistically different from other static classes, and that performance in it correlates with a network's ability to capture temporal information. Thus we use it as a benchmark on current popular networks, which reveals a series of interesting facts, like inflated convolutions bias networks towards classes where motion is not important. We also explore the effect of training on the temporal set, and observe that this leads to better generalization in unseen classes, demonstrating the need for more temporal data. We hope that the proposed dataset of temporal categories will help guide future research in temporal modeling for better video understanding. Laura Sevilla-Lara, Shengxin Zha, Zhicheng Yan 0001, Vedanuj Goswami, Matt Feiszli, Lorenzo Torresani |
WACV | 3 |
| 2020 | Decoupling Representation and Classifier for Long-Tailed Recognition
Bingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan 0001, Albert Gordo, Jiashi Feng, Yannis Kalantidis |
ICLR | 4 |
| 2019 | Graph-Based Global Reasoning NetworksabstractGlobally modeling and reasoning over relations between regions can be beneficial for many computer vision tasks on both images and videos. Convolutional Neural Networks (CNNs) excel at modeling local relations by convolution operations, but they are typically inefficient at capturing global relations between distant regions and require stacking multiple convolution layers. In this work, we propose a new approach for reasoning globally in which a set of features are globally aggregated over the coordinate space and then projected to an interaction space where relational reasoning can be efficiently computed. After reasoning, relation-aware features are distributed back to the original coordinate space for down-stream tasks. We further present a highly efficient instantiation of the proposed approach and introduce the Global Reasoning unit (GloRe unit) that implements the coordinate-interaction space mapping by weighted global pooling and weighted broadcasting, and the relation reasoning via graph convolution on a small graph in interaction space. The proposed GloRe unit is lightweight, end-to-end trainable and can be easily plugged into existing CNNs for a wide range of tasks. Extensive experiments show our GloRe unit can consistently boost the performance of state-of-the-art backbone architectures, including ResNet, ResNeXt, SE-Net and DPN, for both 2D and 3D CNNs, on image classification, semantic segmentation and video action recognition task. Yunpeng Chen, Marcus Rohrbach, Zhicheng Yan 0001, Shuicheng Yan, Jiashi Feng, Yannis Kalantidis |
CVPR | 3 |
| 2019 | DMC-Net: Generating Discriminative Motion Cues for Fast Compressed Video Action RecognitionabstractMotion has shown to be useful for video understanding, where motion is typically represented by optical flow. However, computing flow from video frames is very timeconsuming. Recent works directly leverage the motion vectors and residuals readily available in the compressed video to represent motion at no cost. While this avoids flow computation, it also hurts accuracy since the motion vector is noisy and has substantially reduced resolution, which makes it a less discriminative motion representation. To remedy these issues, we propose a lightweight generator network, which reduces noises in motion vectors and captures fine motion details, achieving a more Discriminative Motion Cue (DMC) representation. Since optical flow is a more accurate motion representation, we train the DMC generator to approximate flow using a reconstruction loss and a generative adversarial loss, jointly with the downstream action classification task. Extensive evaluations on three action recognition benchmarks (HMDB-51, UCF-101, and a subset of Kinetics) confirm the effectiveness of our method. Our full system, consisting of the generator and the classifier, is coined as DMC-Net which obtains high accuracy close to that of using flow and runs two orders of magnitude faster than using optical flow at inference time. Zheng Shou 0001, Xudong Lin 0003, Yannis Kalantidis, Laura Sevilla-Lara, Marcus Rohrbach, Shih-Fu Chang, Zhicheng Yan 0001 |
CVPR | 7 |
| 2019 | Drop an Octave: Reducing Spatial Redundancy in Convolutional Neural Networks With Octave ConvolutionabstractIn natural images, information is conveyed at different frequencies where higher frequencies are usually encoded with fine details and lower frequencies are usually encoded with global structures. Similarly, the output feature maps of a convolution layer can also be seen as a mixture of information at different frequencies. In this work, we propose to factorize the mixed feature maps by their frequencies, and design a novel Octave Convolution (OctConv) operation to store and process feature maps that vary spatially “slower” at a lower spatial resolution reducing both memory and computation cost. Unlike existing multi-scale methods, OctConv is formulated as a single, generic, plug-and-play convolutional unit that can be used as a direct replacement of (vanilla) convolutions without any adjustments in the network architecture. It is also orthogonal and complementary to methods that suggest better topologies or reduce channel-wise redundancy like group or depth-wise convolutions. We experimentally show that by simply replacing convolutions with OctConv, we can consistently boost accuracy for both image and video recognition tasks, while reducing memory and computational cost. An OctConv-equipped ResNet-152 can achieve 82.9% top-1 classification accuracy on ImageNet with merely 22.2 GFLOPs. Yunpeng Chen, Haoqi Fan 0001, Zhicheng Yan 0001, Yannis Kalantidis, Marcus Rohrbach, Shuicheng Yan, Jiashi Feng |
ICCV | 4 |
| 2019 | HACS: Human Action Clips and Segments Dataset for Recognition and Temporal LocalizationabstractThis paper presents a new large-scale dataset for recognition and temporal localization of human actions collected from Web videos. We refer to it as HACS (Human Action Clips and Segments). We leverage consensus and disagreement among visual classifiers to automatically mine candidate short clips from unlabeled videos, which are subsequently validated by human annotators. The resulting dataset is dubbed HACS Clips. Through a separate process we also collect annotations defining action segment boundaries. This resulting dataset is called HACS Segments. Overall, HACS Clips consists of 1.5M annotated clips sampled from 504K untrimmed videos, and HACS Segments contains 139K action segments densely annotated in 50K untrimmed videos spanning 200 action categories. HACS Clips contains more labeled examples than any existing video benchmark. This renders our dataset both a large-scale action recognition benchmark and an excellent source for spatiotemporal feature learning. In our transfer learning experiments on three target datasets, HACS Clips outperforms Kinetics-600, Moments-In-Time and Sports1M as a pretraining source. On HACS Segments, we evaluate state-of-the-art methods of action proposal generation and action localization, and highlight the new challenges posed by our dense temporal annotations. Hang Zhao 0021, Antonio Torralba 0001, Lorenzo Torresani, Zhicheng Yan 0001 |
ICCV | 4 |
| 2019 | Temporal Segment Convolutional Kernel Networks for Sequence Modeling of VideosabstractSequence modeling is crucial for video action recognition. In this paper, we propose temporal segment convolutional kernel networks (TS-CKN), where we take advantage of convolutional neural networks to facilitate the extraction of appearance features, while time sequence is modeled with deep kernel networks. We employ the kernel methods to capture time-varying information of videos and propose a training method for kernel map approximation by matrix backpropagation. This leads to the model named deep kernel networks which can be easily integrated with existing deep learning models such as Resnet. Our approach also samples several video clips sparsely in the video and unifies class predictions from all clips. More importantly, all parameters of our model can be learned by stochastic optimization in an end-to-end manner. We evaluate our method on two standard action recognition datasets including HMDB-51 and UCF-101, achieving the state-of-the-art results. Yanwen Guo 0001, Zhicheng Yan 0001, Jie Guo 0001 |
ICME | 3 |
| 2017 | Exemplar-Based Image and Video Stylization Using Fully Convolutional Semantic FeaturesabstractColor and tone stylization in images and videos strives to enhance unique themes with artistic color and tone adjustments. It has a broad range of applications from professional image postprocessing to photo sharing over social networks. Mainstream photo enhancement softwares, such as Adobe Lightroom and Instagram, provide users with predefined styles, which are often hand-crafted through a trial-and-error process. Such photo adjustment tools lack a semantic understanding of image contents and the resulting global color transform limits the range of artistic styles it can represent. On the other hand, stylistic enhancement needs to apply distinct adjustments to various semantic regions. Such an ability enables a broader range of visual styles. In this paper, we first propose a novel deep learning architecture for exemplar-based image stylization, which learns local enhancement styles from image pairs. Our deep learning architecture consists of fully convolutional networks for automatic semantics-aware feature extraction and fully connected neural layers for adjustment prediction. Image stylization can be efficiently accomplished with a single forward pass through our deep network. To extend our deep network from image stylization to video stylization, we exploit temporal superpixels to facilitate the transfer of artistic styles from image exemplars to videos. Experiments on a number of data sets for image stylization as well as a diverse set of video clips demonstrate the effectiveness of our deep learning architecture. Feida Zhu 0002, Zhicheng Yan 0001, Jiajun Bu, Yizhou Yu |
IEEE Trans. Image Process. | 2 |
| 2016 | Learning Concept Taxonomies from Multi-modal DataabstractWe study the problem of automatically building hypernym taxonomies from textual and visual data.Previous works in taxonomy induction generally ignore the increasingly prominent visual data, which encode important perceptual semantics.Instead, we propose a probabilistic model for taxonomy induction by jointly leveraging text and images.To avoid hand-crafted feature engineering, we design end-to-end features based on distributed representations of images and words.The model is discriminatively trained given a small set of existing ontologies and is capable of building full taxonomies from scratch for a collection of unseen conceptual label items with associated images.We evaluate our model and features on the WordNet hierarchies, where our system outperforms previous approaches by a large gap. Hao Zhang 0025, Zhiting Hu, Yuntian Deng, Mrinmaya Sachan, Zhicheng Yan 0001, Eric P. Xing |
ACL (1) | 5 |
| 2016 | Automatic Photo Adjustment Using Deep Neural NetworksabstractPhoto retouching enables photographers to invoke dramatic visual impressions by artistically enhancing their photos through stylistic color and tone adjustments. However, it is also a time-consuming and challenging task that requires advanced skills beyond the abilities of casual photographers. Using an automated algorithm is an appealing alternative to manual work, but such an algorithm faces many hurdles. Many photographic styles rely on subtle adjustments that depend on the image content and even its semantics. Further, these adjustments are often spatially varying. Existing automatic algorithms are still limited and cover only a subset of these challenges. Recently, deep learning has shown unique abilities to address hard problems. This motivated us to explore the use of deep neural networks (DNNs) in the context of photo editing. In this article, we formulate automatic photo adjustment in a manner suitable for this approach. We also introduce an image descriptor accounting for the local semantics of an image. Our experiments demonstrate that training DNNs using these descriptors successfully capture sophisticated photographic styles. In particular and unlike previous techniques, it can model local adjustments that depend on image semantics. We show that this yields results that are qualitatively and quantitatively better than previous work. Zhicheng Yan 0001, Hao Zhang 0025, Baoyuan Wang, Sylvain Paris, Yizhou Yu |
ACM Trans. Graph. | 1 |
| 2015 | HD-CNN: Hierarchical Deep Convolutional Neural Networks for Large Scale Visual RecognitionabstractIn image classification, visual separability between different object categories is highly uneven, and some categories are more difficult to distinguish than others. Such difficult categories demand more dedicated classifiers. However, existing deep convolutional neural networks (CNN) are trained as flat N-way classifiers, and few efforts have been made to leverage the hierarchical structure of categories. In this paper, we introduce hierarchical deep CNNs (HD-CNNs) by embedding deep CNNs into a two-level category hierarchy. An HD-CNN separates easy classes using a coarse category classifier while distinguishing difficult classes using fine category classifiers. During HDCNN training, component-wise pretraining is followed by global fine-tuning with a multinomial logistic loss regularized by a coarse category consistency term. In addition, conditional executions of fine category classifiers and layer parameter compression make HD-CNNs scalable for largescale visual recognition. We achieve state-of-the-art results on both CIFAR100 and large-scale ImageNet 1000-class benchmark datasets. In our experiments, we build up three different two-level HD-CNNs, and they lower the top-1 error of the standard CNNs by 2:65%, 3:1%, and 1:1%. Zhicheng Yan 0001, Hao Zhang 0025, Robinson Piramuthu, Vignesh Jagadeesh, Dennis DeCoste, Wei Di, Yizhou Yu |
ICCV | 1 |
| 2015 | Multi-facet Learning using Deep Convolutional Neural Network for Person-Related Categories in PhotosabstractThis paper proposes to leverage multiple facets of person photos to improve the training of deep neural networks. Existing studies usually require a lot of labeled images to train deep convolutional networks. Our study suggests exploring multiple datasets and learning effective representation to learn related visual concepts. The practice of learning from multiple facets implicitly enforces to share features for image recognition. We show deep neural network benefits from the learning of multiple person-related categories in photos. Faceted classification systems learn from multiple resources, and alleviate the overfitting problems in deep learning. Moreover, by exploring multiple taxonomies of an object, it provides a finer annotation for the query images. Liangliang Cao, Zhicheng Yan 0001, John R. Smith |
ICMR | 2 |
| 2013 | Sparse similarity matrix learning for visual object retrievalabstractTf-idf weighting scheme is adopted by state-of-the-art object retrieval systems to reflect the difference in discriminability between visual words. However, we argue it is only suboptimal by noting that tf-idf weighting scheme does not take quantization error into account and exploit word correlation. We view tf-idf weights as an example of diagonal Mahalanobis-type similarity matrix and generalize it into a sparse one by selectively activating off-diagonal elements. Our goal is to separate similarity of relevant images from that of irrelevant ones by a safe margin. We satisfy such similarity constraints by learning an optimal similarity metric from labeled data. An effective scheme is developed to collect training data with an emphasis on cases where the tf-idf weights violates the relative relevance constraints. Experimental results on benchmark datasets indicate the learnt similarity metric consistently and significantly outperforms the tf-idf weighting scheme. Zhicheng Yan 0001, Yizhou Yu |
IJCNN | 1 |
| 2013 | Fast nonrigid 3D retrieval using modal space transformabstractNonrigid or deformable 3D objects are common in many application domains. Retrieval of such objects in large databases based on shape similarity is still a challenging problem. In this paper, we first analyze the advantages of functional operators, and further propose a framework to design novel shape signatures for encoding nonrigid object structures. Our approach constructs a context-aware integral kernel operator on a manifold, then applies modal analysis to map this operator into a low-frequency functional representation, called fast functional transform, and finally computes its spectrum as the shape signature. Our method is fast, isometry-invariant, discriminative, and numerically stable with respect to multiple types of perturbations. Jianbo Ye, Zhicheng Yan 0001, Yizhou Yu |
ICMR | 2 |
| 2009 | Context-aware Volume Modeling of Skeletal MusclesabstractAbstract This paper presents an interactive volume modeling method that constructs skeletal muscles from an existing volumetric dataset. Our approach provides users with an intuitive modeling interface and produces compelling results that conform to the characteristic anatomy in the input volume. The algorithmic core of our method is an intuitive anatomy classification approach, suited to accommodate spatial constraints on the muscle volume. The presented work is useful in illustrative visualization, volumetric information fusion and volume illustration that involve muscle modeling, where the spatial context should be faithfully preserved. Zhicheng Yan 0001, Wei Chen 0001, Aidong Lu, David S. Ebert |
Comput. Graph. Forum | 1 |
| 2009 | A Novel Interface for Interactive Exploration of DTI FibersabstractVisual exploration is essential to the visualization and analysis of densely sampled 3D DTI fibers in biological specimens, due to the high geometric, spatial, and anatomical complexity of fiber tracts. Previous methods for DTI fiber visualization use zooming, color-mapping, selection, and abstraction to deliver the characteristics of the fibers. However, these schemes mainly focus on the optimization of visualization in the 3D space where cluttering and occlusion make grasping even a few thousand fibers difficult. This paper introduces a novel interaction method that augments the 3D visualization with a 2D representation containing a low-dimensional embedding of the DTI fibers. This embedding preserves the relationship between the fibers and removes the visual clutter that is inherent in 3D renderings of the fibers. This new interface allows the user to manipulate the DTI fibers as both 3D curves and 2D embedded points and easily compare or validate his or her results in both domains. The implementation of the framework is GPU based to achieve real-time interaction. The framework was applied to several tasks, and the results show that our method reduces the user's workload in recognizing 3D DTI fibers and permits quick and accurate DTI fiber selection. Wei Chen 0001, Zi'ang Ding, Song Zhang 0004, Anna MacKay-Brandt, Stephen Correia, Huamin Qu, John Allen Crow, David F. Tate, Zhicheng Yan 0001, Qunsheng Peng 0001 |
IEEE Trans. Vis. Comput. Graph. | 9 |
| 2009 | Volume Illustration of Muscle from Diffusion Tensor ImagesabstractMedical illustration has demonstrated its effectiveness to depict salient anatomical features while hiding the irrelevant details. Current solutions are ineffective for visualizing fibrous structures such as muscle, because typical datasets (CT or MRI) do not contain directional details. In this paper, we introduce a new muscle illustration approach that leverages diffusion tensor imaging (DTI) data and example-based texture synthesis techniques. Beginning with a volumetric diffusion tensor image, we reformulate it into a scalar field and an auxiliary guidance vector field to represent the structure and orientation of a muscle bundle. A muscle mask derived from the input diffusion tensor image is used to classify the muscle structure. The guidance vector field is further refined to remove noise and clarify structure. To simulate the internal appearance of the muscle, we propose a new two-dimensional example based solid texture synthesis algorithm that builds a solid texture constrained by the guidance vector field. Illustrating the constructed scalar field and solid texture efficiently highlights the global appearance of the muscle as well as the local shape and structure of the muscle fibers in an illustrative fashion. We have applied the proposed approach to five example datasets (four pig hearts and a pig leg), demonstrating plausible illustration and expressiveness. Wei Chen 0001, Zhicheng Yan 0001, Song Zhang 0004, John Allen Crow, David S. Ebert, Ronald M. McLaughlin, Katie B. Mullins, Robert Cooper, Zi'ang Ding |
IEEE Trans. Vis. Comput. Graph. | 2 |