VLDB 2026 Research / reviewers in the wild / expert
Yinpeng Chen
dblp:45/6977
· DBLP profile ↗
58ranked-venue papers
18as first author
36since 2021 · last 2025
0000-0003-1411-225XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 39 · 16 first-author · 18 since 2021Artificial intelligence and machine learning · 36 · 6 first-author · 28 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MELODI: Exploring Memory Compression for Long ContextsabstractWe present MELODI, a novel memory architecture designed to efficiently process long documents using short context windows. The key principle behind MELODI is to represent short-term and long-term memory as a hierarchical compression scheme across both transformer layers and context windows. Specifically, the short-term memory is achieved through recurrent compression of context windows across multiple layers, ensuring smooth transitions between windows. In contrast, the long-term memory performs further compression within a single middle layer and aggregates information across context windows, effectively consolidating crucial information from the entire history. Compared to a strong baseline - the Memorizing Transformer employing dense attention over a large long-term memory (64K key-value pairs) - our method demonstrates superior performance on various long-context datasets while remarkably reducing the memory footprint by a factor of 8. Yinpeng Chen, DeLesley Hutchins, Aren Jansen, Andrey Zhmoginov, David Racz, Jesper Sparre Andersen |
ICLR | 1 |
| 2025 | Exploring Invariance in Images through One-way Wave EquationsabstractIn this paper, we empirically demonstrate that natural images can be reconstructed with high fidelity from compressed representations using a simple first-order norm-plus-linear autoregressive (FINOLA) process—without relying on explicit positional information. Through systematic analysis, we observe that the learned coefficient matrices ($\mathbf{A}$ and $\mathbf{B}$) in FINOLA are typically invertible, and their product, $\mathbf{AB}^{-1}$, is diagonalizable across training runs. This structure enables a striking interpretation: FINOLA’s latent dynamics resemble a system of one-way wave equations evolving in a compressed latent space. Under this framework, each image corresponds to a unique solution of these equations. This offers a new perspective on image invariance, suggesting that the underlying structure of images may be governed by simple, invariant dynamic laws. Our findings shed light on a novel avenue for understanding and modeling visual data through the lens of latent-space dynamics and wave propagation. Yinpeng Chen, Dongdong Chen 0001, Xiyang Dai, Mengchen Liu, Yinan Feng, Youzuo Lin, Lu Yuan 0001, Zicheng Liu 0001 |
ICML | 1 |
| 2025 | Survey of Deep Learning and Physics-Based Approaches in Computational Wave ImagingabstractComputational wave imaging (CWI) extracts hidden structure and physical properties of a volume of material by analyzing wave signals that traverse that volume. Applications include seismic exploration of the Earth's subsurface, acoustic imaging and non-destructive testing in material science, and ultrasound computed tomography in medicine. Current approaches for solving CWI problems can be divided into two categories: those rooted in traditional physics, and those based on deep learning. Physics-based methods stand out for their ability to provide high-resolution and quantitatively accurate estimates of acoustic properties within the medium. However, they can be computationally intensive and are susceptible to ill-posedness and nonconvexity typical of CWI problems. Machine learning-based computational methods have recently emerged, offering a different perspective to address these challenges. Diverse scientific communities have independently pursued the integration of deep learning in CWI. This review discusses how contemporary scientific machine-learning (ML) techniques, and deep neural networks in particular, have been developed to enhance and integrate with traditional physics-based methods for solving CWI problems. We present a structured framework that consolidates existing research spanning multiple domains, including computational imaging, wave physics, and data science. This study concludes with important lessons learned from existing ML-based methods and identifies technical hurdles and emerging trends through a systematic analysis of the extensive literature on this topic. Youzuo Lin, Shihang Feng, James Theiler, Yinpeng Chen, Umberto Villa, Jing Rao, John James Greenhall, Cristian Pantea, Mark A. Anastasio, Brendt Wohlberg |
Proc. IEEE | 4 |
| 2025 | Weak Supervision Multigeophysical Inversion for CO2 Saturation ImagingabstractIn CO2sequestration projects, multi-physics inversion has been widely used to reconstruct various geophysical properties (such as velocity and conductivity). However, the saturation of CO2is not feasible through partial differential equations (PDEs), which poses a challenge to traditional multi-physics inversion techniques in directly inverting CO2saturation from geophysical measurements. Typically, a rock physics model is constructed to tackle this challenge. Nevertheless, the construction of such a model is intricate due to the inherent complexities and uncertainties within subsurface geology and geophysical data. Data-driven inversion methods present an alternative solution, as they can connect CO2saturation to geophysical measurements and directly learn their relationships from labeled data. Nonetheless, the efficacy of these methods hinges on extensive data labeling, incurring considerable costs. To address this challenge, we propose a novel data-driven technique for the inversion of multi-physics data called Weakly Supervised Multi-Geophysical Inversion (WS-MGI), which reduces the need for labeling and promises to be a more cost-effective solution. In particular, we focus on the multi-physics inversion problem from two geophysical data (electromagnetic (EM) and seismic) to CO2saturation. By learning the local relationship between velocity and CO2saturation at a few well logs, we construct the pseudo labels for CO2saturation, thus enabling the data-driven inversion with few labels. We verify our method using synthetic data based on the Kimberlina storage reservoir in California. Experiments show that compared to the supervised counterpart, our method can achieve similar inversion results with only 2% labels. Shihang Feng, Yinpeng Chen, Xitong Zhang, David Alumbaugh, Michael Commer, Youzuo Lin |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Tutorial on Novel Toolkits toward AI for Science on Resource-Constrained Computing SystemsabstractFull Waveform Inversion (FWI) is a technique used to visualize and analyze wave propagation through a medium in order to infer its physical properties. This method relies on computational models and algorithms to simulate and interpret the behavior of waves—such as sound, electromagnetic, or seismic waves—as they travel through different materials. By analyzing how these waves are reflected, refracted, or absorbed by the medium, FWI can provide detailed information about the medium’s internal structure, composition, and physical properties, such as density, elasticity, or internal defects. The traditional process typically involves: 1) Wave Simulation: Using physics-based models to simulate how waves propagate through a medium. This may involve solving complex differential equations that describe wave behavior in different contexts. 2) Data Acquisition: Collecting data on wave interactions with the medium using sensors or other measurement devices. This could include data on wave speed, direction, amplitude, and phase changes. 3) Image Reconstruction: Applying computational techniques, such as inverse problems or tomographic reconstruction, to create images or maps of the medium based on the acquired wave data. 4) Analysis: Interpreting the reconstructed images to deduce the physical properties of the medium. This can involve identifying features like boundaries, interfaces, or anomalies within the medium. Yi Sheng 0001, Junhuan Yang, Hanchen Wang 0003, Yinan Feng, Yinpeng Chen, Youzuo Lin, Weiwen Jiang, Lei Yang 0018 |
CODES+ISSS | 5 |
| 2024 | Efficient Modulation for Vision NetworksabstractIn this work, we present efficient modulation, a novel design for efficient vision networks. We revisit the modulation mechanism, which operates input through convolutional context modeling and feature projection layers, and fuses features via element-wise multiplication and an MLP block. We demonstrate that the abstracted modulation mechanism is particularly well suited for efficient networks and further tailor the modulation design by proposing the efficient modulation (EfficientMod) block, which is considered the essential building block for our networks. Bene- fiting from the prominent representational ability of modulation mechanism and the efficiency of efficient modulation design, our network can accomplish better accuracy-efficiency trade-offs and set new state-of-the-art performance for efficient networks. When integrating EfficientMod block with the vanilla self-attention block, we obtain the hybrid architecture and further improve the performance without sacrificing the efficiency. We carry out comprehensive experiments to verify EfficientMod’s performance. With fewer parameters, our EfficientMod-s performs 0.6 top-1 accuracy better than the prior state-of-the-art approach EfficientFormerV2-s2 without any training tricks and is 25% faster on GPU. Additionally, our method presents a notable improvement in downstream tasks, outperforming EfficientFormerV2-s by 3.6 mIoU on the ADE20K benchmark. Code and checkpoints are available at https://github.com/ma-xu/EfficientMod. Xu Ma 0005, Xiyang Dai, Bin Xiao 0004, Yinpeng Chen, Yun Fu 0001, Lu Yuan 0001 |
ICLR | 5 |
| 2024 | Completing Visual Objects via Bridging Generation and SegmentationabstractThis paper presents a novel approach to object completion, with the primary goal of reconstructing a complete object from its partially visible components. Our method, named MaskComp, delineates the completion process through iterative stages of generation and segmentation. In each iteration, the object mask is provided as an additional condition to boost image generation, and, in return, the generated images can lead to a more accurate mask by fusing the segmentation of images. We demonstrate that the combination of one generation and one segmentation stage effectively functions as a mask denoiser. Through alternation between the generation and segmentation stages, the partial object mask is progressively refined, providing precise shape guidance and yielding superior object completion results. Our experiments demonstrate the superiority of MaskComp over existing approaches, e.g., ControlNet and Stable Diffusion, establishing it as an effective solution for object completion. Xiang Li 0106, Yinpeng Chen, Chung-Ching Lin, Hao Chen 0102, Kai Hu 0010, Rita Singh, Bhiksha Raj, Zicheng Liu 0001 |
ICML | 2 |
| 2024 | Auto-Linear Phenomenon in Subsurface ImagingabstractSubsurface imaging involves solving full waveform inversion (FWI) to predict geophysical properties from measurements. This problem can be reframed as an image-to-image translation, with the usual approach being to train an encoder-decoder network using paired data from two domains: geophysical property and measurement. A recent seminal work (InvLINT) demonstrates there is only a linear mapping between the latent spaces of the two domains, and the decoder requires paired data for training. This paper extends this direction by demonstrating that only linear mapping necessitates paired data, while both the encoder and decoder can be learned from their respective domains through self-supervised learning. This unveils an intriguing phenomenon (named Auto-Linear) where the self-learned features of two separate domains are automatically linearly correlated. Compared with existing methods, our Auto-Linear has four advantages: (a) solving both forward and inverse modeling simultaneously, (b) reducing model size, (c) enhanced performance, especially when the paired data is limited, and (d) strong generalization ability of the trained encoder and decoder. Yinan Feng, Yinpeng Chen, Shihang Feng, Youzuo Lin |
ICML | 2 |
| 2023 | Find Beauty in the Rare: Contrastive Composition Feature Clustering for Nontrivial Cropping Box RegressionabstractAutomatic image cropping algorithms aim to recompose images like human-being photographers by generating the cropping boxes with improved composition quality. Cropping box regression approaches learn the beauty of composition from annotated cropping boxes. However, the bias of annotations leads to quasi-trivial recomposing results, which has an obvious tendency to the average location of training samples. The crux of this predicament is that the task is naively treated as a box regression problem, where rare samples might be dominated by normal samples, and the composition patterns of rare samples are not well exploited. Observing that similar composition patterns tend to be shared by the cropping boundaries annotated nearly, we argue to find the beauty of composition from the rare samples by clustering the samples with similar cropping boundary annotations, i.e., similar composition patterns. We propose a novel Contrastive Composition Clustering (C2C) to regularize the composition features by contrasting dynamically established similar and dissimilar pairs. In this way, common composition patterns of multiple images can be better summarized, which especially benefits the rare samples and endows our model with better generalizability to render nontrivial results. Extensive experimental results show the superiority of our model compared with prior arts. We also illustrate the philosophy of our design with an interesting analytical visualization. Yinpeng Chen, Hao Lu 0003, Zhiguo Cao 0001, Weicai Zhong |
AAAI | 2 |
| 2023 | Detection Hub: Unifying Object Detection Datasets via Query Adaptation on Language EmbeddingabstractCombining multiple datasets enables performance boost on many computer vision tasks. But similar trend has not been witnessed in object detection when combining multiple datasets due to two inconsistencies among detection datasets: taxonomy difference and domain gap. In this paper, we address these challenges by a new design (named Detection Hub) that is dataset-aware and category-aligned. It not only mitigates the dataset inconsistency but also provides coherent guidance for the detector to learn across multiple datasets. In particular, the dataset-aware design is achieved by learning a dataset embedding that is used to adapt object queries as well as convolutional kernels in detection heads. The categories across datasets are semantically aligned into a unified space by replacing one-hot category representations with word embedding and leveraging the semantic coherence of language embedding. Detection Hub fulfills the benefits of large data on object detection. Experiments demonstrate that joint training on multiple datasets achieves significant performance gains over training on each dataset alone. Detection Hub further achieves SoTA performance on UODB benchmark with wide variety of datasets. Lingchen Meng, Xiyang Dai, Yinpeng Chen, Pengchuan Zhang, Dongdong Chen 0001, Mengchen Liu, Zuxuan Wu, Lu Yuan 0001, Yu-Gang Jiang 0001 |
CVPR | 3 |
| 2023 | Masked Video Distillation: Rethinking Masked Feature Modeling for Self-supervised Video Representation LearningabstractBenefiting from masked visual modeling, self-supervised video representation learning has achieved remarkable progress. However, existing methods focus on learning representations from scratch through reconstructing low-level features like raw pixel values. In this paper, we propose masked video distillation (MVD), a simple yet effective two-stage masked feature modeling framework for video representation learning: firstly we pretrain an image (or video) model by recovering low-level features of masked patches, then we use the resulting features as targets for masked feature modeling. For the choice of teacher models, we observe that students taught by video teachers perform better on temporally-heavy video tasks, while image teachers transfer stronger spatial representations for spatially-heavy video tasks. Visualization analysis also indicates different teachers produce different learned patterns for students. To leverage the advantage of different teachers, we design a spatial-temporal co-teaching method for MVD. Specifically, we distill student models from both video teachers and image teachers by masked feature modeling. Extensive experimental results demonstrate that video transformers pre-trained with spatial-temporal co-teaching outperform models distilled with a single teacher on a multitude of video datasets. Our MVD with vanilla ViT achieves state-of-the-art performance compared with previous methods on several challenging video downstream tasks. For example, with the ViT-Large model, our MVD achieves 86.4% and 76.7% Top-1 accuracy on Kinetics-400 and Something-Something-v2, outperforming VideoMAE by 1.2% and 2.4% respectively. When a larger ViT-Huge model is adopted, MVD achieves the state-of-the-art performance with 77.3% Top-1 accuracy on Something-Something-v2. Code will be available at https://github.com/ruiwang2021/mvd. Rui Wang 0095, Dongdong Chen 0001, Zuxuan Wu, Yinpeng Chen, Xiyang Dai, Mengchen Liu, Lu Yuan 0001, Yu-Gang Jiang 0001 |
CVPR | 4 |
| 2023 | Improving Adversarial Robustness of Masked Autoencoders via Test-time Frequency-domain PromptingabstractIn this paper, we investigate the adversarial robustness of vision transformers that are equipped with BERT pretraining (e.g., BEiT, MAE). A surprising observation is that MAE has significantly worse adversarial robustness than other BERT pretraining methods. This observation drives us to rethink the basic differences between these BERT pretraining methods and how these differences affect the robustness against adversarial perturbations. Our empirical analysis reveals that the adversarial robustness of BERT pretraining is highly related to the reconstruction target, i.e., predicting the raw pixels of masked image patches will degrade more adversarial robustness of the model than predicting the semantic context, since it guides the model to concentrate more on medium-/high-frequency components of images. Based on our analysis, we provide a simple yet effective way to boost the adversarial robustness of MAE. The basic idea is using the dataset-extracted domain knowledge to occupy the medium-/high-frequency of images, thus narrowing the optimization space of adversarial perturbations. Specifically, we group the distribution of pretraining data and optimize a set of cluster-specific visual prompts on frequency domain. These prompts are incorporated with input images through prototype-based prompt selection during test period. Extensive evaluation shows that our method clearly boost MAE’s adversarial robustness while maintaining its clean performance on ImageNet-1k classification. Our code is available at: https://github.com/shikiw/RobustMAE. Qidong Huang, Xiaoyi Dong, Dongdong Chen 0001, Yinpeng Chen, Lu Yuan 0001, Gang Hua 0001, Weiming Zhang 0001, Nenghai Yu |
ICCV | 4 |
| 2023 | Zero-Shot Human-Object Interaction (HOI) Classification by Bridging Generative and Contrastive Image-Language ModelsabstractExisting studies in Human-Object Interaction (HOI) classification rely on costly human-annotated labels. The goal of this paper is to study a new zero-shot setup to remove the dependency on ground-truth labels. We propose a novel Heterogenous Teacher-Student (HTS) framework and a new loss function. HTS employs a generative pretrained image captioner as the teacher and a contrastive pre-trained classifier as the student. HTS combines the discriminability from generative pre-training and efficiency from contrastive pre-training. To facilitate learning of HOI in this setup, we introduce pseudo-label filtering which aggregates HOI probabilities from multiple regional captions to supervise the student. To enhance the multi-label learning of the student on few-shot classes, we design LogSumExp (LSE)-Sign loss which features a dynamic gradient re-weighting mechanism. Eventually, the student achieves 49.6 mAP on the HICO dataset without using ground truth, becoming a new state-of-the-art method that outperforms supervised approaches. Code is available. Yinpeng Chen, Jenq-Neng Hwang, Zicheng Liu 0001 |
ICIP | 2 |
| 2023 | Layer Grafted Pre-training: Bridging Contrastive Learning And Masked Image Modeling For Label-Efficient Representations
Ziyu Jiang, Yinpeng Chen, Mengchen Liu, Dongdong Chen 0001, Xiyang Dai, Lu Yuan 0001, Zicheng Liu 0001, Zhangyang Wang |
ICLR | 2 |
| 2023 | EFWI: Multiparameter Benchmark Datasets for Elastic Full Waveform Inversion of Geophysical Properties
Shihang Feng, Hanchen Wang 0003, Chengyuan Deng, Yinan Feng, Yinpeng Chen, Youzuo Lin |
NeurIPS | 8 |
| 2023 | PaintSeg: Painting Pixels for Training-free SegmentationabstractThe paper introduces PaintSeg, a new unsupervised method for segmenting objects without any training. We propose an adversarial masked contrastive painting (AMCP) process, which creates a contrast between the original image and a painted image in which a masked area is painted using off-the-shelf generative models. During the painting process, inpainting and outpainting are alternated, with the former masking the foreground and filling in the background, and the latter masking the background while recovering the missing part of the foreground object. Inpainting and outpainting, also referred to as I-step and O-step, allow our method to gradually advance the target segmentation mask toward the ground truth without supervision or training. PaintSeg can be configured to work with a variety of prompts, e.g. coarse masks, boxes, scribbles, and points. Our experimental results demonstrate that PaintSeg outperforms existing approaches in coarse mask-prompt, box-prompt, and point-prompt segmentation tasks, providing a training-free solution suitable for unsupervised segmentation. Code: https://github.com/lxa9867/PaintSeg. Xiang Li 0106, Chung-Ching Lin, Yinpeng Chen, Zicheng Liu 0001, Jinglu Wang, Rita Singh, Bhiksha Raj |
NeurIPS | 3 |
| 2023 | Learning from Rich Semantics and Coarse Locations for Long-tailed Object DetectionabstractLong-tailed object detection (LTOD) aims to handle the extreme data imbalance in real-world datasets, where many tail classes have scarce instances. One popular strategy is to explore extra data with image-level labels, yet it produces limited results due to (1) semantic ambiguity---an image-level label only captures a salient part of the image, ignoring the remaining rich semantics within the image; and (2) location sensitivity---the label highly depends on the locations and crops of the original image, which may change after data transformations like random cropping.
To remedy this, we propose RichSem, a simple but effective method, which is robust to learn rich semantics from coarse locations without the need of accurate bounding boxes. RichSem leverages rich semantics from images, which are then served as additional ``soft supervision'' for training detectors. Specifically, we add a semantic branch
to our detector to learn these soft semantics and enhance feature representations for long-tailed object detection. The semantic branch is only used for training and is removed during inference. RichSem achieves consistent improvements on both overall and rare-category of LVIS under different backbones and detectors.
Our method achieves state-of-the-art performance without requiring complex training and testing procedures. Moreover, we show the effectiveness of our method on other long-tailed datasets with additional experiments. Lingchen Meng, Xiyang Dai, Dongdong Chen 0001, Yinpeng Chen, Mengchen Liu, Zuxuan Wu, Lu Yuan 0001, Yu-Gang Jiang 0001 |
NeurIPS | 5 |
| 2022 | Mobile-Former: Bridging MobileNet and TransformerabstractWe present Mobile-Former, a parallel design of MobileNet and transformer with a two-way bridge in between. This structure leverages the advantages of MobileNet at local processing and transformer at global interaction. And the bridge enables bidirectional fusion of local and global features. Different from recent works on vision transformer, the transformer in Mobile-Former contains very few tokens (e.g. 6 or fewer tokens) that are randomly initialized to learn global priors, resulting in low computational cost. Combining with the proposed light-weight cross attention to model the bridge, Mobile-Former is not only computationally efficient, but also has more representation power. It outperforms MobileNetV3 at low FLOP regime from 25M to 500M FLOPs on ImageNet classification. For instance, Mobile-Former achieves 77.9% top-1 accuracy at 294M FLOPs, gaining 1.3% over MobileNetV3 but saving 17% of computations. When transferring to object detection, Mobile-Former outperforms MobileNetV3 by 8.6 AP in RetinaNet framework. Furthermore, we build an efficient end-to-end detector by replacing backbone, encoder and decoder in DETR with Mobile-Former, which outperforms DETR by 1.3 AP but saves 52% of computational cost and 36% of parameters. Code will be released at https://github.com/aaboys/mobileformer. Yinpeng Chen, Xiyang Dai, Dongdong Chen 0001, Mengchen Liu, Xiaoyi Dong, Lu Yuan 0001, Zicheng Liu 0001 |
CVPR | 1 |
| 2022 | Reduce Information Loss in Transformers for Pluralistic Image InpaintingabstractTransformers have achieved great success in pluralistic image inpainting recently. However, we find existing transformer based solutions regard each pixel as a token, thus suffer from information loss issue from two aspects: 1) They downsample the input image into much lower resolutions for efficiency consideration, incurring information loss and extra misalignment for the boundaries of masked regions. 2) They quantize 2563RGB pixels to a small number (such as 512) of quantized pixels. The indices of quantized pixels are used as tokens for the inputs and prediction targets of transformer. Although an extra CNN network is used to upsample and refine the low-resolution results, it is difficult to retrieve the lost information back. To keep input information as much as possible, we propose a new transformer based framework “PUT”. Specifically, to avoid input downsampling while maintaining the computation efficiency, we design a patch-based auto-encoder P-VQVAE, where the encoder converts the masked image into non-overlapped patch tokens and the decoder recovers the masked regions from the inpainted tokens while keeping the unmasked regions unchanged. To eliminate the information loss caused by quantization, an Un-Quantized Transformer (UQ-Transformer) is applied, which directly takes the features from P-VQVAE encoder as input without quantization and regards the quantized tokens only as prediction targets. Extensive experiments show that PUT greatly outperforms state-of-the-art methods on image fidelity, especially for large masked regions and complex large-scale datasets. Qiankun Liu 0001, Zhentao Tan, Dongdong Chen 0001, Qi Chu 0001, Xiyang Dai, Yinpeng Chen, Mengchen Liu, Lu Yuan 0001, Nenghai Yu |
CVPR | 6 |
| 2022 | BEVT: BERT Pretraining of Video TransformersabstractThis paper studies the BERT pretraining of video transformers. It is a straightforward but worth-studying extension given the recent success from BERT pretraining of image transformers. We introduce BEVT which decouples video representation learning into spatial representation learning and temporal dynamics learning. In particular, BEVT first performs masked image modeling on image data, and then conducts masked image modeling jointly with masked video modeling on video data. This design is motivated by two observations: 1) transformers learned on image datasets provide decent spatial priors that can ease the learning of video transformers, which are often times computationally-intensive if trained from scratch; 2) discriminative clues, i.e., spatial and temporal information, needed to make correct predictions vary among different videos due to large intra-class and inter-class variations. We conduct extensive experiments on three challenging video benchmarks where BEVT achieves very promising results. On Kinetics 400, for which recognition mostly relies on discriminative spatial representations, BEVT achieves comparable results to strong supervised baselines. On Something-Something-V2 and Diving 48, which contain videos relying on temporal dynamics, BEVT outperforms by clear margins all alternative baselines and achieves state-of-the-art performance with a 71.4% and 87.2% Top-1 accuracy respectively. Code is available at https://github.com/xyzforever/BEVT. Rui Wang 0095, Dongdong Chen 0001, Zuxuan Wu, Yinpeng Chen, Xiyang Dai, Mengchen Liu, Yu-Gang Jiang 0001, Luowei Zhou, Lu Yuan 0001 |
CVPR | 4 |
| 2022 | Should All Proposals Be Treated Equally in Object Detection?
Yunsheng Li, Yinpeng Chen, Xiyang Dai, Dongdong Chen 0001, Mengchen Liu, Pei Yu, Lu Yuan 0001, Zicheng Liu 0001, Nuno Vasconcelos |
ECCV (25) | 2 |
| 2022 | Hyprogan: Breaking the Dimensional wall From Human to AnimeabstractImage translation from human faces to anime ones brings a low-end, efficient way to create animation characters for animation industry. However, due to the significant inter-domain difference between anime images and human photos, existing image-to-image translation approaches cannot address this task well. To solve this dilemma, we propose HyProGAN, an exemplar-guided image-to-image translation model without paired data. The key contribution of HyPro-GAN is that it introduces a novel hybrid and progressive training strategy that expands the unidirectional translation between two domains into the bidirectional intra-domain and inter-domain translation. To enhance the consistency between input and output, we further propose a local masking loss to align the facial features between the human face and the generated anime face. Extensive experiments demonstrate the superiority of HyProGAN against state-of-the-art models. Yinpeng Chen, Zhiguo Cao 0001, Hao Lu 0003, Weicai Zhong |
ICIP | 1 |
| 2022 | Unsupervised Learning of Full-Waveform Inversion: Connecting CNN and Partial Differential Equation in a Loop
Xitong Zhang, Yinpeng Chen, Sharon X. Huang, Zicheng Liu 0001, Youzuo Lin |
ICLR | 3 |
| 2022 | An Intriguing Property of Geophysics InversionabstractInversion techniques are widely used to reconstruct subsurface physical properties (e.g., velocity, conductivity) from surface-based geophysical measurements (e.g., seismic, electric/magnetic (EM) data). The problems are governed by partial differential equations (PDEs) like the wave or Maxwell’s equations. Solving geophysical inversion problems is challenging due to the ill-posedness and high computational cost. To alleviate those issues, recent studies leverage deep neural networks to learn the inversion mappings from measurements to the property directly. In this paper, we show that such a mapping can be well modeled by a very shallow (but not wide) network with only five layers. This is achieved based on our new finding of an intriguing property: a near-linear relationship between the input and output, after applying integral transform in high dimensional space. In particular, when dealing with the inversion from seismic data to subsurface velocity governed by a wave equation, the integral results of velocity with Gaussian kernels are linearly correlated to the integral of seismic data with sine kernels. Furthermore, this property can be easily turned into a light-weight encoder-decoder network for inversion. The encoder contains the integration of seismic data and the linear transformation without need for fine-tuning. The decoder only consists of a single transformer block to reverse the integral of velocity. Experiments show that this interesting property holds for two geophysics inversion problems over four different datasets. Compared to much deeper InversionNet, our method achieves comparable accuracy, but consumes significantly fewer parameters Yinan Feng, Yinpeng Chen, Shihang Feng, Zicheng Liu 0001, Youzuo Lin |
ICML | 2 |
| 2022 | Design What You Desire: Icon Generation from Orthogonal Application and Theme LabelsabstractGenerative adversarial networks,(GANs) have been trained to be professional artists able to create stunning artworks such as face generation and image style transfer. In this paper, we focus on a realistic business scenario: automated generation of customizable icons given desired mobile applications and theme styles. We first introduce a theme-application icon dataset, termed AppIcon, where each icon has two orthogonal theme and app labels. By investigating a strong baseline StyleGAN2, we observe mode collapse caused by the entanglement of the orthogonal labels. To solve this challenge, we propose IconGAN composed of a conditional generator and dual discriminators with orthogonal augmentations, and a contrastive feature disentanglement strategy is further designed to regularize the feature space of the two discriminators. Compared with other approaches, IconGAN indicates a superior advantage on the AppIcon benchmark. Further analysis also justifies the effectiveness of disentangling app and theme representations. Our project will be released at: https://github.com/architect-road/IconGAN. Yinpeng Chen, Min Shi 0004, Hao Lu 0003, Zhiguo Cao 0001, Weicai Zhong |
ACM Multimedia | 1 |
| 2022 | OpenFWI: Large-scale Multi-structural Benchmark Datasets for Full Waveform InversionabstractFull waveform inversion (FWI) is widely used in geophysics to reconstruct high-resolution velocity maps from seismic data. The recent success of data-driven FWI methods results in a rapidly increasing demand for open datasets to serve the geophysics community. We present OpenFWI, a collection of large-scale multi-structural benchmark datasets, to facilitate diversified, rigorous, and reproducible research on FWI. In particular, OpenFWI consists of $12$ datasets ($2.1$TB in total) synthesized from multiple sources. It encompasses diverse domains in geophysics (interface, fault, CO$_2$ reservoir, etc.), covers different geological subsurface structures (flat, curve, etc.), and contain various amounts of data samples (2K - 67K). It also includes a dataset for 3D FWI. Moreover, we use OpenFWI to perform benchmarking over four deep learning methods, covering both supervised and unsupervised learning regimes. Along with the benchmarks, we implement additional experiments, including physics-driven methods, complexity analysis, generalization study, uncertainty quantification, and so on, to sharpen our understanding of datasets and methods. The studies either provide valuable insights into the datasets and the performance, or uncover their current limitations. We hope OpenFWI supports prospective research on FWI and inspires future open-source efforts on AI for science. All datasets and related information can be accessed through our website at https://openfwi-lanl.github.io/ Chengyuan Deng, Shihang Feng, Hanchen Wang 0003, Xitong Zhang, Yinan Feng, Qili Zeng, Yinpeng Chen, Youzuo Lin |
NeurIPS | 8 |
| 2022 | SC-UDA: Style and Content Gaps aware Unsupervised Domain Adaptation for Object DetectionabstractCurrent state-of-the-art object detectors can have significant performance drop when deployed in the wild due to domain gaps with training data. Unsupervised Domain Adaptation (UDA) is a promising approach to adapt detectors for new domains/environments without any expensive label cost. Previous mainstream UDA works for object detection usually focused on image-level and/or feature-level adaptation by using adversarial learning methods. In this work, we show that such adversarial-based methods can only reduce domain style gap, but cannot address the domain content gap that is also important for object detectors. To overcome this limitation, we propose the SC-UDA framework to concurrently reduce both gaps: We propose fine-grained domain style transfer to reduce the style gaps with finer image details preserved for detecting small objects; Then we leverage the pseudo label-based self-training to reduce content gaps; To address pseudo label error accumulation during self-training, novel optimizations are proposed, including uncertainty-based pseudo labeling and imbalanced mini-batch sampling strategy. Experiment results show that our approach consistently outperforms prior state-of-the-art methods (up to 8.6%, 2.7% and 2.5% mAP on three UDA benchmarks). Fuxun Yu, Di Wang 0003, Yinpeng Chen, Nikolaos Karianakis, Pei Yu, Dimitrios Lymberopoulos, Sidi Lu, Weisong Shi, Xiang Chen 0010 |
WACV | 3 |
| 2022 | Class-attribute inconsistency learning for novelty detection
Shuaiyuan Du, Chaoyi Hong, Yinpeng Chen, Zhiguo Cao 0001 |
Pattern Recognit. | 3 |
| 2021 | Dynamic Head: Unifying Object Detection Heads With AttentionsabstractThe complex nature of combining localization and classification in object detection has resulted in the flourished development of methods. Previous works tried to improve the performance in various object detection heads but failed to present a unified view. In this paper, we present a novel dynamic head framework to unify object detection heads with attentions. By coherently combining multiple self-attention mechanisms between feature levels for scale-awareness, among spatial locations for spatial-awareness, and within output channels for task-awareness, the proposed approach significantly improves the representation ability of object detection heads without any computational overhead. Further experiments demonstrate that the effectiveness and efficiency of the proposed dynamic head on the COCO benchmark. With a standard ResNeXt-101-DCN backbone, we largely improve the performance over popular object detectors and achieve a new state-of-the-art at 54.0 AP. The code will be released at https://github.com/microsoft/DynamicHead. Xiyang Dai, Yinpeng Chen, Bin Xiao 0004, Dongdong Chen 0001, Mengchen Liu, Lu Yuan 0001, Lei Zhang 0001 |
CVPR | 2 |
| 2021 | Dynamic Transfer for Multi-Source Domain AdaptationabstractRecent works of multi-source domain adaptation focus on learning a domain-agnostic model, of which the parameters are static. However, such a static model is difficult to handle conflicts across multiple domains, and suffers from a performance degradation in both source domains and target domain. In this paper, we present dynamic transfer to address domain conflicts, where the model parameters are adapted to samples. The key insight is that adapting model across domains is achieved via adapting model across samples. Thus, it breaks down source domain barriers and turns multi-source domains into a single-source domain. This also simplifies the alignment between source and target domains, as it only requires the target domain to be aligned with any part of the union of source domains. Furthermore, we find dynamic transfer can be simply modeled by aggregating residual matrices and a static convolution matrix. Experimental results show that, without using domain labels, our dynamic transfer outperforms the state-of-the-art method by more than 3% on the large multi-source domain adaptation datasets – DomainNet. Source code is at https://github.com/liyunsheng13/DRT. Yunsheng Li, Lu Yuan 0001, Yinpeng Chen, Nuno Vasconcelos |
CVPR | 3 |
| 2021 | Dynamic DETR: End-to-End Object Detection with Dynamic AttentionabstractIn this paper, we present a novel Dynamic DETR (Detection with Transformers) approach by introducing dynamic attentions into both the encoder and decoder stages of DETR to break its two limitations on small feature resolution and slow training convergence. To address the first limitation, which is due to the quadratic computational complexity of the self-attention module in Transformer encoders, we propose a dynamic encoder to approximate the Transformer encoder’s attention mechanism using a convolution-based dynamic encoder with various attention types. Such an encoder can dynamically adjust attentions based on multiple factors such as scale importance, spatial importance, and representation (i.e., feature dimension) importance. To mitigate the second limitation of learning difficulty, we introduce a dynamic decoder by replacing the cross-attention module with a ROI-based dynamic attention in the Transformer decoder. Such a decoder effectively assists Transformers to focus on region of interests from a coarse-to-fine manner and dramatically lowers the learning difficulty, leading to a much faster convergence with fewer training epochs. We conduct a series of experiments to demonstrate our advantages. Our Dynamic DETR significantly reduces the training epochs (by 14×), yet results in a much better performance (by 3.6 on mAP). Meanwhile, in the standard 1× setup with ResNet-50 backbone, we archive a new state-of-the-art performance that further proves the learning effectiveness of the proposed approach. Xiyang Dai, Yinpeng Chen, Pengchuan Zhang, Lu Yuan 0001, Lei Zhang 0001 |
ICCV | 2 |
| 2021 | Improve Unsupervised Pretraining for Few-label TransferabstractUnsupervised pretraining has achieved great success and many recent works have shown unsupervised pretraining can achieve comparable or even slightly better transfer performance than supervised pretraining on downstream target datasets. But in this paper, we find this conclusion may not hold when the target dataset has very few labeled samples for finetuning, i.e., few-label transfer. We analyze the possible reason from the clustering perspective: 1) The clustering quality of target samples is of great importance to few-label transfer; 2) Though contrastive learning is essential to learn how to cluster, its clustering quality is still inferior to supervised pretraining due to lack of label supervision. Based on the analysis, we interestingly discover that only involving some unlabeled target domain into the unsupervised pretraining can improve the clustering quality, subsequently reducing the transfer performance gap with supervised pretraining. This finding also motivates us to propose a new progressive few-label transfer algorithm for real applications, which aims to maximize the transfer performance under a limited annotation budget. To support our analysis and proposed method, we conduct extensive experiments on nine different target datasets. Experimental results show our proposed method can significantly boost the few-label transfer performance of unsupervised pretraining. Suichan Li, Dongdong Chen 0001, Yinpeng Chen, Lu Yuan 0001, Lei Zhang 0001, Qi Chu 0001, Bin Liu 0016, Nenghai Yu |
ICCV | 3 |
| 2021 | MicroNet: Improving Image Recognition with Extremely Low FLOPsabstractThis paper aims at addressing the problem of substantial performance degradation at extremely low computational cost (e.g. 5M FLOPs on ImageNet classification). We found that two factors, sparse connectivity and dynamic activation function, are effective to improve the accuracy. The former avoids the significant reduction of network width, while the latter mitigates the detriment of reduction in network depth. Technically, we propose micro-factorized convolution, which factorizes a convolution matrix into low rank matrices, to integrate sparse connectivity into convolution. We also present a new dynamic activation function, named Dynamic Shift Max, to improve the non-linearity via maxing out multiple dynamic fusions between an input feature map and its circular channel shift. Building upon these two new operators, we arrive at a family of networks, named MicroNet, that achieves significant performance gains over the state of the art in the low FLOP regime. For instance, under the constraint of 12M FLOPs, MicroNet achieves 59.4% top-1 accuracy on ImageNet classification, outperforming MobileNetV3 by 9.6%. Source code is at https://github.com/liyunsheng13/micronet. Yunsheng Li, Yinpeng Chen, Xiyang Dai, Dongdong Chen 0001, Mengchen Liu, Lu Yuan 0001, Zicheng Liu 0001, Lei Zhang 0001, Nuno Vasconcelos |
ICCV | 2 |
| 2021 | Revisiting Dynamic Convolution via Matrix Decomposition
Yunsheng Li, Yinpeng Chen, Xiyang Dai, Mengchen Liu, Dongdong Chen 0001, Lu Yuan 0001, Zicheng Liu 0001, Nuno Vasconcelos |
ICLR | 2 |
| 2021 | Stronger NAS with Weaker PredictorsabstractNeural Architecture Search (NAS) often trains and evaluates a large number of architectures. Recent predictor-based NAS approaches attempt to alleviate such heavy computation costs with two key steps: sampling some architecture-performance pairs and fitting a proxy accuracy predictor. Given limited samples, these predictors, however, are far from accurate to locate top architectures due to the difficulty of fitting the huge search space. This paper reflects on a simple yet crucial question: if our final goal is to find the best architecture, do we really need to model the whole space well?. We propose a paradigm shift from fitting the whole architecture space using one strong predictor, to progressively fitting a search path towards the high-performance sub-space through a set of weaker predictors. As a key property of the weak predictors, their probabilities of sampling better architectures keep increasing. Hence we only sample a few well-performed architectures guided by the previously learned predictor and estimate a new better weak predictor. This embarrassingly easy framework, dubbed WeakNAS, produces coarse-to-fine iteration to gradually refine the ranking of sampling space. Extensive experiments demonstrate that WeakNAS costs fewer samples to find top-performance architectures on NAS-Bench-101 and NAS-Bench-201. Compared to state-of-the-art (SOTA) predictor-based NAS methods, WeakNAS outperforms all with notable margins, e.g., requiring at least 7.5x less samples to find global optimal on NAS-Bench-101. WeakNAS can also absorb their ideas to boost performance more. Further, WeakNAS strikes the new SOTA result of 81.3% in the ImageNet MobileNet Search Space. The code is available at: https://github.com/VITA-Group/WeakNAS. Xiyang Dai, Dongdong Chen 0001, Yinpeng Chen, Mengchen Liu, Zhangyang Wang, Zicheng Liu 0001, Lu Yuan 0001 |
NeurIPS | 4 |
| 2021 | Cross-Domain Complementary Learning Using Pose for Multi-Person Part SegmentationabstractSupervised deep learning with pixel-wise training labels has great successes on multi-person part segmentation. However, data labeling at pixel-level is very expensive. To solve the problem, people have been exploring to use synthetic data to avoid the data labeling. Although it is easy to generate labels for synthetic data, the results are much worse compared to those using real data and manual labeling. The degradation of the performance is mainly due to the domain gap, i.e., the discrepancy of the pixel value statistics between real and synthetic data. In this paper, we observe that real and synthetic humans both have a skeleton (pose) representation. We found that the skeletons can effectively bridge the synthetic and real domains during the training. Our proposed approach takes advantage of the rich and realistic variations of the real data and the easily obtainable labels of the synthetic data to learn multi-person part segmentation on real images without any human-annotated labels. Through experiments, we show that without any human labeling, our method performs comparably to several state-of-the-art approaches which require human labeling on Pascal-Person-Parts and COCO-DensePose datasets. On the other hand, if part labels are also available in the real-images during training, our method outperforms the supervised state-of-the-art methods by a large margin. We further demonstrate the generalizability of our method on predicting novel keypoints in real images where no real data labels are available for the novel keypoints detection. Code and pre-trained models are available at https://github.com/kevinlin311tw/CDCL-human-part-segmentation. Yinpeng Chen, Zicheng Liu 0001, Ming-Ting Sun |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | CPTNet: Cascade Pose Transform Network for Single Image Talking Head Animation
Ke Xian, Yinpeng Chen, Zhiguo Cao 0001, Weicai Zhong |
ACCV (4) | 4 |
| 2020 | Rethinking Classification and Localization for Object DetectionabstractTwo head structures (i.e. fully connected head and convolution head) have been widely used in R-CNN based detectors for classification and localization tasks. However, there is a lack of understanding of how does these two head structures work for these two tasks. To address this issue, we perform a thorough analysis and find an interesting fact that the two head structures have opposite preferences towards the two tasks. Specifically, the fully connected head (fc-head) is more suitable for the classification task, while the convolution head (conv-head) is more suitable for the localization task. Furthermore, we examine the output feature maps of both heads and find that fc-head has more spatial sensitivity than conv-head. Thus, fc-head has more capability to distinguish a complete object from part of an object, but is not robust to regress the whole object. Based upon these findings, we propose a Double-Head method, which has a fully connected head focusing on classification and a convolution head for bounding box regression. Without bells and whistles, our method gains +3.5 and +2.8 AP on MS COCO dataset from Feature Pyramid Network (FPN) baselines with ResNet-50 and ResNet-101 backbones, respectively. Yue Wu 0008, Yinpeng Chen, Lu Yuan 0001, Zicheng Liu 0001, Yun Fu 0001 |
CVPR | 2 |
| 2020 | Dynamic Convolution: Attention Over Convolution KernelsabstractLight-weight convolutional neural networks (CNNs) suffer performance degradation as their low computational budgets constrain both the depth (number of convolution layers) and the width (number of channels) of CNNs, resulting in limited representation capability. To address this issue, we present Dynamic Convolution, a new design that increases model complexity without increasing the network depth or width. Instead of using a single convolution kernel per layer, dynamic convolution aggregates multiple parallel convolution kernels dynamically based upon their attentions, which are input dependent. Assembling multiple kernels is not only computationally efficient due to the small kernel size, but also has more representation power since these kernels are aggregated in a non-linear way via attention. By simply using dynamic convolution for the state-of-the-art architecture MobileNetV3-Small, the top-1 accuracy of ImageNet classification is boosted by 2.9% with only 4% additional FLOPs and 2.9 AP gain is achieved on COCO keypoint detection. Yinpeng Chen, Xiyang Dai, Mengchen Liu, Dongdong Chen 0001, Lu Yuan 0001, Zicheng Liu 0001 |
CVPR | 1 |
| 2020 | Dynamic ReLU
Yinpeng Chen, Xiyang Dai, Mengchen Liu, Dongdong Chen 0001, Lu Yuan 0001, Zicheng Liu 0001 |
ECCV (19) | 1 |
| 2020 | DA-NAS: Data Adapted Pruning for Efficient Neural Architecture Search
Xiyang Dai, Dongdong Chen 0001, Mengchen Liu, Yinpeng Chen, Lu Yuan 0001 |
ECCV (27) | 4 |
| 2019 | Large Scale Incremental LearningabstractModern machine learning suffers from \textit{catastrophic forgetting} when learning new classes incrementally. The performance dramatically degrades due to the missing data of old classes. Incremental learning methods have been proposed to retain the knowledge acquired from the old classes, by using knowledge distilling and keeping a few exemplars from the old classes. However, these methods struggle to \textbf{scale up to a large number of classes}. We believe this is because of the combination of two factors: (a) the data imbalance between the old and new classes, and (b) the increasing number of visually similar classes. Distinguishing between an increasing number of visually similar classes is particularly challenging, when the training data is unbalanced. We propose a simple and effective method to address this data imbalance issue. We found that the last fully connected layer has a strong bias towards the new classes, and this bias can be corrected by a linear model. With two bias parameters, our method performs remarkably well on two large datasets: ImageNet (1000 classes) and MS-Celeb-1M (10000 classes), outperforming the state-of-the-art algorithms by 11.1\% and 13.2\% respectively. Yue Wu 0008, Yinpeng Chen, Yuancheng Ye, Zicheng Liu 0001, Yandong Guo, Yun Fu 0001 |
CVPR | 2 |
| 2018 | Reinforced Temporal Attention and Split-Rate Transfer for Depth-Based Person Re-identification
Nikolaos Karianakis, Zicheng Liu 0001, Yinpeng Chen, Stefano Soatto |
ECCV (5) | 3 |
| 2015 | ImmerseBoard: Immersive Telepresence Experience using a Digital WhiteboardabstractImmerseBoard is a system for remote collaboration through a digital whiteboard that gives participants a 3D immersive experience, enabled only by an RGBD camera (Microsoft Kinect) mounted on the side of a large touch display. Using 3D processing of the depth images, life-sized rendering, and novel visualizations, ImmerseBoard emulates writing side-by-side on a physical whiteboard, or alternatively on a mirror. User studies involving three tasks show that compared to standard video conferencing with a digital whiteboard, ImmerseBoard provides participants with a quantitatively better ability to estimate their remote partners' eye gaze direction, gesture direction, intention, and level of agreement. Moreover, these quantitative capabilities translate qualitatively into a heightened sense of being together and a more enjoyable experience. ImmerseBoard's form factor is suitable for practical and easy installation in homes and offices. Keita Higuchi, Yinpeng Chen, Philip A. Chou, Zhengyou Zhang, Zicheng Liu 0001 |
CHI | 2 |
| 2015 | VTouch: Vision-enhanced interaction for large touch displaysabstractWe propose a system that augments touch input with visual understanding of the user to improve interaction with a large touch-sensitive display. A commodity color plus depth sensor such as Microsoft Kinect adds the visual modality and enables new interactions beyond touch. Through visual analysis, the system understands where the user is, who the user is, and what the user is doing even before the user touches the display. Such information is used to enhance interaction in multiple ways. For example, a user can use simple gestures to bring up menu items such as color palette and soft keyboard; menu items can be shown where the user is and can follow the user; hovering can show information to the user before the user commits to touch; the user can perform different functions (for example writing and erasing) with different hands; and the user's preference profile can be maintained, distinct from other users. User studies are conducted and the users very much appreciate the value of these and other enhanced interactions. Yinpeng Chen, Zicheng Liu 0001, Philip A. Chou, Zhengyou Zhang |
ICME | 1 |
| 2014 | Compression of human body sequences using graph Wavelet Filter BanksabstractThe next step in immersive communication beyond video from a single camera is object-based free viewpoint video, which is the capture and compression of a dynamic object such that it can be reconstructed and viewed from an arbitrary viewpoint. The moving human body is a particularly useful subclass of dynamic object for object-based free viewpoint video relevant to both telepresence and entertainment. In this paper, we compress moving human body sequences by applying recently developed Graph Wavelet Filter Banks to time-varying geometry and color signals living on a mesh representation of the human body. This model-based approach significantly outperforms state-of-the-art coding of the human body represented as ordinary depth plus color video sequences. Ha Q. Nguyen 0001, Philip A. Chou, Yinpeng Chen |
ICASSP | 3 |
| 2013 | Tensor-Based Human Body ModelingabstractIn this paper, we present a novel approach to model 3D human body with variations on both human shape and pose, by exploring a tensor decomposition technique. 3D human body modeling is important for 3D reconstruction and animation of realistic human body, which can be widely used in Tele-presence and video game applications. It is challenging due to a wide range of shape variations over different people and poses. The existing SCAPE model is popular in computer vision for modeling 3D human body. However, it considers shape and pose deformations separately, which is not accurate since pose deformation is person-dependent. Our tensor-based model addresses this issue by jointly modeling shape and pose deformations. Experimental results demonstrate that our tensor-based model outperforms the SCAPE model quite significantly. We also apply our model to capture human body using Microsoft Kinect sensors with excellent results. Yinpeng Chen, Zicheng Liu 0001, Zhengyou Zhang |
CVPR | 1 |
| 2011 | A home-based adaptive mixed reality rehabilitation systemabstractThis paper presents an interactive home-based adaptive mixed reality system (HAMRR) for upper extremity stroke rehabilitation. This home-based system is an extension of a previously designed and currently implemented clinical system. The goal of HAMRR is to restore motor function to chronic stroke survivors by providing an engaging long-term reaching task therapy at home. The HAMMR system tracks movement of the wrist and torso, and provides real-time, post-trial, and post-set multimodal feedback to encourage the stroke survivor to self-assess his or her movement and engage in active learning of new movement strategies. This experiential media system uses a computational adaptation scheme to create a continuously challenging and unique multi-year therapy experience through the use of multiple, integrated audio and visual feedback streams. Novel design features include creating an over-arching story for the participant, the ability of the system to adapt the feedback over multiple time scales, and the ability for this system to integrate into any home. Diana Siwiak, Nicole Lehrer, Michael Baran, Yinpeng Chen, Margaret Duff, Todd Ingalls, Thanassis Rikakis |
ACM Multimedia | 4 |
| 2010 | Adaptive mixed reality stroke rehabilitation: system architecture and evaluation metricsabstractThis paper presents a novel system architecture and evaluation metrics for an Adaptive Mixed Reality Rehabilitation (AMRR) system for stroke patient. This system provides a purposeful, engaging, hybrid (visual, auditory and physical) scene that encourages patients to improve their performance of a reaching and grasping task and promotes learning of generalizable movement strategies. This system is adaptive in that it provides assistive adaptation tools to help the rehabilitation team customize the training strategy. Our key insight is to combine the patients, rehabilitation team, multimodal hybrid environments and adaptation tools together as an adaptive experiential mixed reality system. Yinpeng Chen, Nicole Lehrer, Hari Sundaram, Thanassis Rikakis |
MMSys | 1 |
| 2008 | A dynamic decision network framework for online media adaptation in stroke rehabilitationabstractIn this article, we present a media adaptation framework for an immersive biofeedback system for stroke patient rehabilitation. In our biofeedback system, media adaptation refers to changes in audio/visual feedback as well as changes in physical environment. Effective media adaptation frameworks help patients recover generative plans for arm movement with potential for significantly shortened therapeutic time. The media adaptation problem has significant challenges—(a) high dimensionality of adaptation parameter space; (b) variability in the patient performance across and within sessions; (c) the actual rehabilitation plan is typically a non-first-order Markov process, making the learning task hard. Our key insight is to understand media adaptation as a real-time feedback control problem. We use a mixture-of-experts based Dynamic Decision Network (DDN) for online media adaptation. We train DDN mixtures per patient, per session. The mixture models address two basic questions—(a) given a specific adaptation suggested by the domain experts, predict the patient performance, and (b) given the expected performance, determine the optimal adaptation decision. The questions are answered through an optimality criterion based search on DDN models trained in previous sessions. We have also developed new validation metrics and have very good results for both questions on actual stroke rehabilitation data. Yinpeng Chen, Hari Sundaram, Thanassis Rikakis, Sheng-Min Liu |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2007 | Media adaptation framework in biofeedback system for stroke patient rehabilitationabstractIn this paper, we present a media adaptation framework for an immersive biofeedback system for stroke patient rehabilitation. In our biofeedback system, media adaptation refers to changes in audio/visual feedback as well as changes in physical environment. Effective media adaptation frameworks help patients recover generative plans for arm movement with potential for significantly shortened therapeutic time. The media adaptation problem has significant challenges - (a) high dimensionality of adaptation parameter space (b) variability in the patient performance across and within sessions(c) the actual rehabilitation plan is typically a non first-order Markov process, making the learning task hard. Yinpeng Chen, Hari Sundaram, Thanassis Rikakis, Sheng-Min Liu |
ACM Multimedia | 1 |
| 2007 | A Computational Estimate of the Physical Effort in Human Poses
Yinpeng Chen, Hari Sundaram, Jodi James |
MMM (2) | 1 |
| 2006 | Basis Projection for Linear Transform Approximation in Real-Time ApplicationsabstractThis paper aims to develop a novel framework to systematically trade-off computational complexity with output distortion, in linear multimedia transforms, in an optimal manner. The problem is important in real-time systems where the computational resources available are time-dependent. We solve the real-time adaptation problem by developing an approximate transform framework. There are three key contributions of this paper – (a) a fast basis approximation framework that allows us to store signal independent partial transform results to be used in real-time, (b) estimating the complexity distortion curve for the linear transform using a basis set and (c) determining optimal operating points and a meta-data embedding algorithm for images that allows for real-time adaptation. We have applied this approach on the FFT transform with excellent results. Yinpeng Chen, Hari Sundaram |
ICASSP (2) | 1 |
| 2006 | The design of a real-time, multimodal biofeedback system for stroke patient rehabilitationabstractThis paper presents a novel real-time, multi-modal biofeedback system for stroke patient therapy. The problem is important as traditional mechanisms of rehabilitation are monotonous, and do not incorporate detailed quantitative assessment of recovery in addition to traditional clinical schemes. We have been working on developing an experiential media system that integrates task dependent physical therapy and cognitive stimuli within an interactive, multimodal environment. The environment provides a purposeful, engaging, visual and auditory scene in which patients can practice functional therapeutic reaching tasks, while receiving different types of simultaneous feedback indicating measures of both performance and results. There are three contributions of this paper - (a) identification of features and goals for the functional task (b) The development of sophisticated feedback (auditory and visual) mechanisms that match the semantics of action of the task. We additionally develop novel action-feedback coupling mechanisms. (c) New metrics to validate the ability of the system to promote learnability, stylization and engagement. We have validated the system for nine subjects with excellent results. Yinpeng Chen, He Huang 0002, Richard Isaac Wallis, Hari Sundaram, Thanassis Rikakis, Todd Ingalls, Loren Olson, Jiping He |
ACM Multimedia | 1 |
| 2006 | A real-time, multimodal biofeedback system for stroke patient rehabilitationabstractThis paper presents a novel real-time, multi-modal biofeedback system for stoke patient therapy. The problem is important as traditional mechanisms of rehabilitation are monotonous, and do not incorporate detailed quantitative assessment of recovery in addition to traditional clinical schemes. We have been working on developing an experiential media system that integrates task dependent physical therapy and cognitive stimuli within an interactive, multimodal environment. The environment provides a purposeful, engaging, visual and auditory scene in which patients can practice functional therapeutic reaching tasks, while receiving different types of simultaneous feedback indicating measures of both performance and results. There are two contributions of this paper - (a) identification of features and goals for the functional task, (b) the development of sophisticated feedback (auditory and visual) mechanisms that match the semantics of action of the task. Yinpeng Chen, Richard Isaac Wallis, Hari Sundaram, Thanassis Rikakis, Todd Ingalls, Loren Olson, Jiping He |
ACM Multimedia | 1 |
| 2005 | A computationally efficient 3D shape rejection algorithmabstractIn this paper, we present an efficient 3D shape rejection algorithm for unlabeled 3D markers. The problem is important in domains such as rehabilitation and the performing arts. There are three key innovations in our approach-(a) a multi-resolution shape representation using Haar wavelets for unlabeled markers, (b) a multi-resolution shape metric and (c) a shape rejection algorithm that is predicated on the simple idea that we do not need to compute the entire distance to conclude that two shapes are dissimilar. We tested the approach on a real-world pose classification problem with excellent results. We achieved a classification accuracy of 98% with an order of magnitude improvement in terms of computational complexity over a baseline shape matching algorithm. Yinpeng Chen, Hari Sundaram |
ICME | 1 |
| 2005 | Media processing workflow design and execution with ARIAabstractRecently, we introduced a novel ARchitecture for Interactive Arts (ARIA) middleware that processes, filters, and fuses sensory inputs and actuates responses in real-time while providing various Quality of Service (QoS) guarantees. The objective of ARIA is to incorporate realtime, sensed, and archived media and audience responses into live performances, on demand. An ARIA media workflow graph describes how the data sensed through media capture devices will be processed and what audio-visual responses will be actuated. Thus, each data object streamed between ARIA processing components is subject to transformations, as described by a media workflow graph. The media capture and processing components, such as media filters and fusion operators, are programmable and adaptable; i.e, the delay, size, frequency, and quality/precision characteristics of individual operators can be controlled via a number of parameters. In [1, 4, 5], we developed static and dynamic optimization algorithms which maximize the quality of the actuated responses, minimize the corresponding delay and the resource usage. In this demonstration, we present the ARIA GUI and the underlying kernel. More specifically, we describe how to design a media processing workflow, with adaptive operators, using the ARIA GUI and how to use the various optimization and adaptation alternatives provided by the ARIA kernel to execute media processing workflows. Lina Peng, Gisik Kwon, K. Selçuk Candan, Kyung Dong Ryu, Karam S. Chatha, Hari Sundaram, Yinpeng Chen |
ACM Multimedia | 7 |
| 2005 | Estimating Complexity of 2D ShapesabstractThis paper deals with the problem of estimating 2D shape complexity. This has important applications in computer vision as well as in developing efficient shape classification algorithms. We define shape complexity using correlates of Kolmogorov complexity-entropy measures of global distance and local angle, and a measure of shape randomness. We tested our algorithm on synthetic and real world datasets with excellent results. We also conducted user studies that indicate that our measure is highly correlated with human perception. They also reveal an intuitive shape sensitivity curve-simple shapes are easily distinguished by small complexity variations, while complex shapes require significant complexity differences to be differentiated Yinpeng Chen, Hari Sundaram |
MMSP | 1 |