VLDB 2026 Research / reviewers in the wild / expert
Stella X. Yu
dblp:58/5089
· DBLP profile ↗
87ranked-venue papers
13as first author
38since 2021 · last 2025
0000-0002-3507-5761ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 72 · 13 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 62 · 7 first-author · 23 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Systems, architecture and hardware · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Let Humanoids Hike! Integrative Skill Development on Complex TrailsabstractHiking on complex trails demands balance, agility, and adaptive decision-making over unpredictable terrain. Current humanoid research remains fragmented and inadequate for hiking: locomotion focuses on motor skills without long-term goals or situational awareness, while semantic navigation overlooks real-world embodiment and local terrain variability. We propose training humanoids to hike on complex trails, fostering integrative skill development across visual perception, decision making, and motor execution.We develop LEGO-H, a learning framework that enables a humanoid with vision to hike complex trails independently. It has two key innovations. 1) A Temporal Vision Transformer anticipates future steps to guide locomotion, unifying local movement and goal-directed navigation. 2) Latent representations of joint movement patterns combined with hierarchical metric learning allow smooth policy transfer from privileged training to real-world training. These techniques enable LEGO-H to handle diverse physical and environmental challenges without relying on predefined motion patterns. Experiments on diverse simulated hiking trails and humanoids with different morphologies demonstrate LEGO-H’s robustness and versatility, establishing a strong foundation for future humanoid development. Kwan-Yee Lin, Stella X. Yu |
CVPR | 2 |
| 2025 | Open Ad-hoc Categorization with Contextualized Feature LearningabstractAdaptive categorization of visual scenes is essential for AI agents to handle changing tasks. Unlike fixed common categories for plants or animals, ad-hoc categories, such as things to sell at a garage sale, are created dynamically to achieve specific tasks. We study open ad-hoc categorization, where the goal is to infer novel concepts and categorize images based on a given context, a small set of labeled exemplars, and some unlabeled data.We have two key insights: 1) recognizing ad-hoc categories relies on the same perceptual processes as common categories; 2) novel concepts can be discovered semantically by expanding contextual cues or visually by clustering similar patterns. We propose OAK, a simple model that introduces a single learnable context token into CLIP, trained with CLIP’s objective of aligning visual and textual features and GCD’s objective of clustering similar images.On Stanford and Clevr-4 datasets, OAK consistently achieves the state-of-art in accuracy and concept discovery across multiple categorizations, including 87.4% novel accuracy on Stanford Mood, surpassing CLIP and GCD by over 50%. Moreover, OAK generates interpretable saliency maps, focusing on hands for Action, faces for Mood, and backgrounds for Location, promoting transparency and trust while enabling accurate and flexible categorization. Zilin Wang 0009, Sangwoo Mo, Stella X. Yu, Sima Behpour, Liu Ren 0001 |
CVPR | 3 |
| 2025 | Visually Consistent Hierarchical Image ClassificationabstractHierarchical classification predicts labels across multiple levels of a taxonomy, e.g., from coarse-level \textit{Bird} to mid-level \textit{Hummingbird} to fine-level \textit{Green hermit}, allowing flexible recognition under varying visual conditions. It is commonly framed as multiple single-level tasks, but each level may rely on different visual cues. Distinguishing \textit{Bird} from \textit{Plant} relies on {\it global features} like {\it feathers} or {\it leaves}, while separating \textit{Anna's hummingbird} from \textit{Green hermit} requires {\it local details} such as {\it head coloration}.
Prior methods improve accuracy using external semantic supervision, but such statistical learning criteria fail to ensure consistent visual grounding at test time, resulting in incorrect hierarchical classification. We propose, for the first time, to enforce \textit{internal visual consistency} by aligning fine-to-coarse predictions through intra-image segmentation. Our method outperforms zero-shot CLIP and state-of-the-art baselines on hierarchical classification benchmarks, achieving both higher accuracy and more consistent predictions. It also improves internal image segmentation without requiring pixel-level annotations. Seulki Park, Youren Zhang, Stella X. Yu, Sara Beery, Jonathan Huang |
ICLR | 3 |
| 2025 | Test-Time Canonicalization by Foundation Models for Robust PerceptionabstractPerception in the real world requires robustness to diverse viewing conditions. Existing approaches often rely on specialized architectures or training with predefined data augmentations, limiting adaptability. Taking inspiration from mental rotation in human vision, we propose FoCal, a test-time robustness framework that transforms the input into the most typical view. At inference time, FoCal explores a set of transformed images and chooses the one with the highest likelihood under foundation model priors. This test-time optimization boosts robustness while requiring no retraining or architectural changes. Applied to models like CLIP and SAM, it significantly boosts robustness across a wide range of transformations, including 2D and 3D rotations, contrast and lighting shifts, and day-night changes. We also explore potential applications in active vision. By reframing invariance as a test-time optimization problem, FoCal offers a general and scalable approach to robustness. Our code is available at: https://github.com/sutkarsh/focal . Utkarsh Singhal, Ryan Feng, Stella X. Yu, Atul Prakash 0001 |
ICML | 3 |
| 2025 | Normalize Filters! Classical Wisdom for Deep VisionabstractClassical image filters, such as those for averaging or differencing, are carefully normalized to ensure consistency, interpretability, and to avoid artifacts like intensity shifts, halos, or ringing. In contrast, convolutional filters learned end-to-end in deep networks lack such constraints. Although they may resemble wavelets and blob/edge detectors, they are not normalized in the same or any way. Consequently, when images undergo atmospheric transfer, their responses become distorted, leading to incorrect outcomes. We address this limitation by proposing filter normalization, followed by learnable scaling and shifting, akin to batch normalization. This simple yet effective modification ensures that the filters are atmosphere-equivariant, enabling co-domain symmetry. By integrating classical filtering principles into deep learning (applicable to both convolutional neural networks and convolution-dependent vision transformers), our method achieves significant improvements on artificial and natural intensity variation benchmarks. Our ResNet34 could even outperform CLIP by a large margin. Our analysis reveals that unnormalized filters degrade performance, whereas filter normalization regularizes learning, promotes diversity, and improves robustness and generalization. Gustavo Pérez, Stella X. Yu |
NeurIPS | 2 |
| 2025 | Novel View Synthesis from A Few Glimpses via Test-Time Natural Video CompletionabstractGiven just a few glimpses of a scene, can you imagine the movie playing out as the camera glides through it? That’s the lens we take on sparse-input novel view synthesis, not only as filling spatial gaps between widely spaced views, but also as completing a natural video unfolding through space. We recast the task as test-time natural video completion, using powerful priors from pretrained video diffusion models to hallucinate plausible in-between views. Our zero-shot, generation-guided framework produces pseudo views at novel camera
poses, modulated by an uncertainty-aware mechanism for spatial coherence. These synthesized frames densify supervision for 3D Gaussian Splatting (3D-GS) for scene reconstruction, especially in under-observed regions. An iterative feedback loop lets 3D geometry and 2D view synthesis inform each other, improving both the scene reconstruction and the generated views. The result is coherent, high-fidelity renderings from sparse inputs without any scene-specific training or fine-tuning. On LLFF, DTU, DL3DV, and MipNeRF-360, our method significantly outperforms strong 3D-GS baselines under extreme sparsity. Our project page is at https://decayale.github.io/project/SV2CGS. Yixing Wang, Stella X. Yu |
NeurIPS | 3 |
| 2024 | Unsupervised Feature Learning with Emergent Data-Driven PrototypicalityabstractGiven a set of images, our goal is to map each image to a point in a feature space such that, not only point proximity indicates visual similarity, but where it is located directly encodes how prototypical the image is according to the dataset. Our key insight is to perform unsupervised feature learning in hyperbolic instead of Euclidean space, where the distance between points still reflects image similarity, yet we gain additional capacity for representing prototypicality with the location of the point: The closer it is to the origin, the more prototypical it is. The latter property is simply emergent from optimizing the metric learning objective: The image similar to many training instances is best placed at the center of corresponding points in Euclidean space, but closer to the origin in hyperbolic space. We propose an unsupervised feature learning algorithm in Hyperbolic space with sphere pACKing. HACK first generates uniformly packed particles in the Poincaré ball of hyperbolic space and then assigns each image uniquely to a particle. With our feature mapper simply trained to spread out training instances in hyperbolic space, we observe that images move closer to the origin with congealing - a warping process that aligns all the images and makes them appear more common and similar to each other, validating our idea of unsupervised prototypicality discovery. We demonstrate that our data-driven prototypicality provides an easy and superior unsupervised instance selection to reduce sample complexity, increase model generalization with atypical instances and robustness with typical ones. Yunhui Guo, Youren Zhang, Yubei Chen, Stella X. Yu |
CVPR | 4 |
| 2024 | Pose-Aware Self-supervised Learning with Viewpoint Trajectory Regularization
Yubei Chen, Stella X. Yu |
ECCV (21) | 3 |
| 2024 | Learning Hierarchical Image Segmentation For Recognition and By RecognitionabstractLarge vision and language models learned directly through image-text associations often lack detailed visual substantiation, whereas image segmentation tasks are treated separately from recognition, supervisedly learned without interconnections.
Our key observation is that, while an image can be recognized in multiple ways, each has a consistent part-and-whole visual organization. Segmentation thus should be treated not as an end task to be mastered through supervised learning, but as an internal process that evolves with and supports the ultimate goal of recognition.
We propose to integrate a hierarchical segmenter into the recognition process,
{\it train} and {\it adapt} the entire model solely on image-level recognition objectives. We learn hierarchical segmentation {\it for free} alongside recognition, automatically uncovering part-to-whole relationships that not only underpin but also enhance recognition.
Enhancing the Vision Transformer (ViT) with adaptive segment tokens and graph pooling, our model surpasses ViT in unsupervised part-whole discovery, semantic segmentation, image classification, and efficiency. Notably, our model (trained on {\it unlabeled} 1M ImageNet images) outperforms SAM (trained on 11M images and 1 billion masks) by absolute 8\% in mIoU on PartImageNet object segmentation. Tsung-Wei Ke, Sangwoo Mo, Stella X. Yu |
ICLR | 3 |
| 2024 | SkinCON: Towards Consensus for the Uncertainty of Skin Cancer Sub-typing Through Distribution Regularized Adaptive Predictive Sets (DRAPS)
Zhihang Ren, Xinrong Xie, Erik P. Duhaime, Kathy Fang, Tapabrata Chakraborti, Yunhui Guo, Stella X. Yu, David Whitney |
MICCAI (1) | 9 |
| 2024 | Zero-shot Building Attribute Extraction from Large-Scale Vision and Language ModelsabstractExisting building recognition methods, exemplified by BRAILS, utilize supervised learning to extract information from satellite and street-view images for classification and segmentation. However, each task module requires human-annotated data, hindering the scalability and robustness to regional variations and annotation imbalances. In response, we propose a new zero-shot workflow for building attribute extraction that utilizes large-scale vision and language models to mitigate reliance on external annotations. The proposed workflow contains two key components: image-level captioning and segment-level captioning for the building images based on the vocabularies pertinent to structural and civil engineering. These two components generate descriptive captions by computing feature representations of the image and the vocabularies, and facilitating a semantic match between the visual and textual representations. Consequently, our framework offers a promising avenue to enhance AI-driven captioning for building attribute extraction in the structural and civil engineering domains, ultimately reducing reliance on human annotations while bolstering performance and adaptability. Sangryul Jeon, Frank McKenna, Stella X. Yu |
WACV | 5 |
| 2024 | VEATIC: Video-based Emotion and Affect Tracking in Context DatasetabstractHuman affect recognition has been a significant topic in psychophysics and computer vision. However, the currently published datasets have many limitations. For example, most datasets contain frames that contain only information about facial expressions. Due to the limitations of previous datasets, it is very hard to either understand the mechanisms for affect recognition of humans or generalize well on common cases for computer vision models trained on those datasets. In this work, we introduce a brand new large dataset, the Video-based Emotion and Affect Tracking in Context Dataset (VEATIC), that can conquer the limitations of the previous datasets. VEATIC has 124 video clips from Hollywood movies, documentaries, and home videos with continuous valence and arousal ratings of each frame via real-time annotation. Along with the dataset, we propose a new computer vision task to infer the affect of the selected character via both context and character information in each video frame. Additionally, we propose a simple model to benchmark this new computer vision task. We also compare the performance of the pretrained model using our dataset with other similar datasets. Experiments show the competing results of our pretrained model via VEATIC, indicating the generalizability of VEATIC. Our dataset is available at https://veatic.github.io. Zhihang Ren, Jefferson Ortega, Yunhui Guo, Stella X. Yu, David Whitney |
WACV | 6 |
| 2024 | Open Long-Tailed Recognition in a Dynamic WorldabstractReal world data often exhibits a long-tailed and open-ended (i.e., with unseen classes) distribution. A practical recognition system must balance between majority (head) and minority (tail) classes, generalize across the distribution, and acknowledge novelty upon the instances of unseen classes (open classes). We define Open Long-Tailed Recognition++ (OLTR++) as learning from such naturally distributed data and optimizing for the classification accuracy over a balanced test set which includes both known and open classes. OLTR++ handles imbalanced classification, few-shot learning, open-set recognition, and active learning in one integrated algorithm, whereas existing classification approaches often focus only on one or two aspects and deliver poorly over the entire spectrum. The key challenges are: 1) how to share visual knowledge between head and tail classes, 2) how to reduce confusion between tail and open classes, and 3) how to actively explore open classes with learned knowledge. Our algorithm, OLTR++, maps images to a feature space such that visual concepts can relate to each other through a memory association mechanism and a learned metric (dynamic meta-embedding) that both respects the closed world classification of seen classes and acknowledges the novelty of open classes. Additionally, we propose an active learning scheme based on visual memory, which learns to recognize open classes in a data-efficient manner for future expansions. On three large-scale open long-tailed datasets we curated from ImageNet (object-centric), Places (scene-centric), and MS1M (face-centric) data, as well as three standard benchmarks (CIFAR-10-LT, CIFAR-100-LT, and iNaturalist-18), our approach, as a unified framework, consistently demonstrates competitive performance. Notably, our approach also shows strong potential for the active exploration of open classes and the fairness analysis of minority groups. Ziwei Liu 0002, Zhongqi Miao, Xiaohang Zhan, Boqing Gong, Stella X. Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Cut and Learn for Unsupervised Object Detection and Instance SegmentationabstractWe propose Cut-and-LEaRn (CutLER), a simple approach for training unsupervised object detection and seg-mentation models. We leverage the property of self-supervised models to ‘discover’ objects without supervision and amplify it to train a state-of-the-art localization model without any human labels. CutLER first uses our proposed MaskCut approach to generate coarse masks for multiple objects in an image, and then learns a detector on these masks using our robust loss function. We further improve performance by self-training the model on its predictions. Compared to prior work, CutLER is simpler, compatible with different detection architectures, and detects multiple objects. CutLER is also a zero-shot unsupervised detector and improves detection performance AP50by over 2.7× on 11 benchmarks across domains like video frames, paintings, sketches, etc. With finetuning, CutLER serves as a low-shot detector surpassing MoCo-v2 by 7.3% Apboxand 6.6% Apmaskon COCO when training with 5% labels. Xudong Wang 0007, Rohit Girdhar, Stella X. Yu, Ishan Misra |
CVPR | 3 |
| 2023 | Bootstrapping Objectness from Videos by Relaxed Common Fate and Visual GroupingabstractWe study learning object segmentation from unlabeled videos. Humans can easily segment moving objects without knowing what they are. The Gestalt law of common fate, i.e., what move at the same speed belong together, has inspired unsupervised object discovery based on motion segmentation. However, common fate is not a reliable indicator of objectness: Parts of an articulated / deformable object may not move at the same speed, whereas shadows / reflections of an object always move with it but are not part of it. Our insight is to bootstrap objectness by first learning image features from relaxed common fate and then refining them based on visual appearance grouping within the image itself and across images statistically. Specifically, we learn an image segmenter first in the loop of approximating optical flow with constant segment flow plus small within-segment residual flow, and then by refining it for more coherent appearance and statistical figure-ground relevance. On unsupervised video object segmentation, using only ResNet and convolutional heads, our model surpasses the state-of-the-art by absolute gains of 7/9/5% on DAVIS16 / STv2 / FBMS59 respectively, demonstrating the effectiveness of our ideas. Our code is publicly available. Long Lian, Zhirong Wu, Stella X. Yu |
CVPR | 3 |
| 2023 | Learning to Transform for Generalizable Instance-wise InvarianceabstractComputer vision research has long aimed to build systems that are robust to transformations found in natural data. Traditionally, this is done using data augmentation or hard-coding invariances into the architecture. However, too much or too little invariance can hurt, and the correct amount is unknown a priori and dependent on the instance. Ideally, the appropriate invariance would be learned from data and inferred at test-time.We treat invariance as a prediction problem. Given any image, we predict a distribution over transformations can and average over them to make invariant predictions. Combined with a graphical model approach, this distribution forms a flexible, generalizable, and adaptive form of invariance. Our experiments show that it can be used to align datasets and discover prototypes, adapt to out-of-distribution poses, and generalize invariance across classes. When as data augmentation, our method shows accuracy and robustness gains on CIFAR 10, CIFAR10-LT, and TinyImageNet. Utkarsh Singhal, Carlos Esteves, Ameesh Makadia, Stella X. Yu |
ICCV | 4 |
| 2023 | The Audio-Visual BatVision Dataset for Research on Sight and SoundabstractVision research showed remarkable success in understanding our world, propelled by datasets of images and videos. Sensor data from radar, LiDAR and cameras supports research in robotics and autonomous driving for at least a decade. However, while visual sensors may fail in some conditions, sound has recently shown potential to complement sensor data. Simulated room impulse responses (RIR) in 3D apartment-models became a benchmark dataset for the community, fostering a range of audiovisual research. In simulation, depth is predictable from sound, by learning bat-like perception with a neural network. Concurrently, the same was achieved in reality by using RGB-D images and echoes of chirping sounds. Biomimicking bat perception is an exciting new direction but needs dedicated datasets to explore the potential. Therefore, we collected the BatVision dataset to provide large-scale echoes in complex real-world scenes to the community. We equipped a robot with a speaker to emit chirps and a binaural microphone to record their echoes. Synchronized RGB-D images from the same perspective provide visual labels of traversed spaces. We sampled modern US office spaces to historic French university grounds, indoor and outdoor with large architectural variety. This dataset will allow research on robot echolocation, general audio-visual tasks and sound phænomena unavailable in simulated data. We show promising results for audio-only depth prediction and show how state-of-the-art work developed for simulated data can also succeed on our dataset. Project page: https://amandinebtto.github.io/Batvision-Dataset/ Amandine Brunetto, Sascha Hornauer, Stella X. Yu, Fabien Moutarde |
IROS | 3 |
| 2023 | ResoNet: Noise-Trained Physics-Informed MRI Off-Resonance CorrectionabstractMagnetic Resonance Imaging (MRI) is a powerful medical imaging modality that offers diagnostic information without harmful ionizing radiation. Unlike optical imaging, MRI sequentially samples the spatial Fourier domain (k-space) of the image.
Measurements are collected in multiple shots, or readouts, and in each shot, data along a smooth trajectory is sampled.
Conventional MRI data acquisition relies on sampling k-space row-by-row in short intervals, which is slow and inefficient. More efficient, non-Cartesian sampling trajectories (e.g., Spirals) use longer data readout intervals, but are more susceptible to magnetic field inhomogeneities, leading to off-resonance artifacts. Spiral trajectories cause off-resonance blurring in the image, and the mathematics of this blurring resembles that of optical blurring, where magnetic field variation corresponds to depth and readout duration to aperture size. Off-resonance blurring is a system issue with a physics-based, accurate forward model. We present a physics-informed deep learning framework for off-resonance correction in MRI, which is trained exclusively on synthetic, noise-like data with representative marginal statistics. Our approach allows for fat/water separation and is compatible with parallel imaging acceleration. Through end-to-end training using synthetic randomized data (i.e., noise-like images, coil sensitivities, field maps), we train the network to reverse off-resonance effects across diverse anatomies and contrasts without retraining. We demonstrate the effectiveness of our approach through results on phantom and in-vivo data. This work has the potential to facilitate the clinical adoption of non-Cartesian sampling trajectories, enabling efficient, rapid, and motion-robust MRI scans. Code is publicly available at: https://github.com/mikgroup/ResoNet. Alfredo De Goyeneche, Shreya Ramachandran, Ke Wang 0067, Ekin Karasan, Joseph Y. Cheng, Stella X. Yu, Michael Lustig |
NeurIPS | 6 |
| 2023 | Compact and Optimal Deep Learning with Recurrent Parameter GeneratorsabstractDeep learning has achieved tremendous success by training increasingly large models, which are then compressed for practical deployment. We propose a drastically different approach to compact and optimal deep learning: We decouple the Degrees of freedom (DoF) and the actual number of parameters of a model, optimize a small DoF with predefined random linear constraints for a large model of an arbitrary architecture, in one-stage end-to-end learning.Specifically, we create a recurrent parameter generator (RPG), which repeatedly fetches parameters from a ring and unpacks them onto a large model with random permutation and sign flipping to promote parameter decorrelation. We show that gradient descent can automatically find the best model under constraints with in fact faster convergence.Our extensive experimentation reveals a log-linear relationship between model DoF and accuracy. Our RPG demonstrates remarkable DoF reduction, and can be further pruned and quantized for additional run-time performance gain. For example, in terms of top-1 accuracy on ImageNet, RPG achieves 96% of ResNet18’s performance with only 18% DoF (the equivalent of one convolutional layer) and 52% of ResNet34’s performance with only 0.25% DoF! Our work shows significant potential of constrained neural opti-mization in compact and optimal deep learning. Yubei Chen, Stella X. Yu, Brian Cheung, Yann LeCun |
WACV | 3 |
| 2023 | Local pseudo-attributes for long-tailed recognition
Dong-Jin Kim 0003, Tsung-Wei Ke, Stella X. Yu |
Pattern Recognit. Lett. | 3 |
| 2023 | Modeling Semantic Correlation and Hierarchy for Real-World Wildlife RecognitionabstractWe explore the challenges of human-in-the-loop frameworks to label wildlife recognition datasets with a neural network. In wildlife imagery, the main challenges for a model to assist human annotation are two-fold: (1) the training dataset is usually imbalanced, which makes the model's suggestion biased, and (2) there are complex taxonomies in the classes. We establish a simple and efficient baseline, including the debiasing loss function and the hyperbolic network architecture, to address these issues. Moreover, we propose leveraging the semantic correlation to train the model more effectively by adding a co-occurrence layer to our model during training. We demonstrate the efficacy of our method in both a real-world wildlife areal survey recognition dataset and the public image classification dataset, CIFAR100-LT, CIFAR10-LT, and iNaturalist. Dong-Jin Kim 0003, Zhongqi Miao, Yunhui Guo, Stella X. Yu |
IEEE Signal Process. Lett. | 4 |
| 2022 | CO-SNE: Dimensionality Reduction and Visualization for Hyperbolic DataabstractHyperbolic space can naturally embed hierarchies that often exist in real-world data and semantics. While high-dimensional hyperbolic embeddings lead to better representations, most hyperbolic models utilize low-dimensional embeddings, due to non-trivial optimization and visualization of high-dimensional hyperbolic data. We propose CO-SNE, which extends the Euclidean space visualization tool, t-SNE, to hyperbolic space. Like t-SNE, it converts distances between data points to joint probabilities and tries to minimize the Kullback-Leibler divergence between the joint probabilities of high-dimensional data$X$and low-dimensional embedding$Y$. However, unlike Euclidean space, hyperbolic space is inhomogeneous: A volume could contain a lot more points at a location far from the origin. CO-SNE thus uses hyperbolic normal distributions for$X$and hyperbolic Cauchy instead of t-SNE's Student's t-distribution for$Y$, and it additionally seeks to preserve$X$'s individual distances to the Origin in$Y$. We apply CO-SNE to naturally hyperbolic data and supervisedly learned hyperbolic features. Our results demonstrate that CO-SNE deflates high-dimensional hyperbolic data into a low-dimensional space without losing their hyperbolic characteristics, significantly outperforming popular visualization tools such as PCA, t-SNE, UMAP, and HoroPCA which is also designed for hyperbolic data. Yunhui Guo, Haoran Guo, Stella X. Yu |
CVPR | 3 |
| 2022 | Clipped Hyperbolic Classifiers Are Super-Hyperbolic ClassifiersabstractHyperbolic space can naturally embed hierarchies, unlike Euclidean space. Hyperbolic Neural Networks (HNNs) exploit such representational power by lifting Euclidean features into hyperbolic space for classification, outperforming Euclidean neural networks (ENNs) on datasets with known semantic hierarchies. However, HNNs underperform ENNs on standard benchmarks without clear hierarchies, greatly restricting HNNs' applicability in practice. Our key insight is that HNNs' poorer general classification performance results from vanishing gradients during backpropagation, caused by their hybrid architecture connecting Euclidean features to a hyperbolic classifier. We propose an effective solution by simply clipping the Euclidean feature magnitude while training HNNs. Our experiments demonstrate that clipped HNNs become super-hyperbolic classifiers: They are not only consistently better than HNNs which already outperform ENNs on hierarchical data, but also on-par with ENNs on MNIST, CIFAR10, CIFAR100 and ImageNet benchmarks, with better adversarial robustness and out-of-distribution detection. Yunhui Guo, Xudong Wang 0007, Yubei Chen, Stella X. Yu |
CVPR | 4 |
| 2022 | Unsupervised Hierarchical Semantic Segmentation with Multiview Cosegmentation and Clustering TransformersabstractUnsupervised semantic segmentation aims to discover groupings within and across images that capture object-and view-invariance of a category without external supervision. Grouping naturally has levels of granularity, creating ambiguity in unsupervised segmentation. Existing methods avoid this ambiguity and treat it as a factor outside modeling, whereas we embrace it and desire hierarchical grouping consistency for unsupervised segmentation. We approach unsupervised segmentation as a pixel-wise feature learning problem. Our idea is that a good representation shall reveal not just a particular level of grouping, but any level of grouping in a consistent and predictable manner. We enforce spatial consistency of grouping and bootstrap feature learning with co-segmentation among multiple views of the same image, and enforce semantic consistency across the grouping hierarchy with clustering transformers between coarse- and fine-grained features. We deliver the first data-driven unsupervised hierarchical semantic segmentation method called Hierarchical Segment Grouping (HSG). Capturing visual similarity and statistical co-occurrences, HSG also outperforms existing un-supervised segmentation methods by a large margin on five major object- and scene-centric benchmarks. Tsung-Wei Ke, Jyh-Jing Hwang, Yunhui Guo, Xudong Wang 0007, Stella X. Yu |
CVPR | 5 |
| 2022 | Co-domain Symmetry for Complex-Valued Deep LearningabstractWe study complex-valued scaling as a type of symmetry natural and unique to complex-valued measurements and representations. Deep Complex Networks (DCN) extend real-valued algebra to the complex domain without addressing complex-valued scaling. SurReal extends manifold learning to the complex plane, achieving scaling invariance with manifold distances that discard phase information. Treating complex-valued scaling as a co-domain transformation, we design novel equivariant/invariant layer functions and architectures that exploit co-domain symmetry. We also propose novel complex-valued representations of RGB images, where complex-valued scaling indicates hue shift or correlated changes across color channels. Benchmarked on MSTAR, CIFAR10, CIFAR100, and SVHN, our co-domain symmetric (CDS) classifiers deliver higher accuracy, better generalization, more robustness to co-domain transformations, and lower model bias and variance than DCN and SurReal with far fewer parameters. Utkarsh Singhal, Yifei Xing 0001, Stella X. Yu |
CVPR | 3 |
| 2022 | Debiased Learning from Naturally Imbalanced Pseudo-LabelsabstractPseudo-labels are confident predictions made on unlabeled target data by a classifier trained on labeled source data. They are widely used for adapting a model to unlabeled data, e.g., in a semi-supervised learning setting. Our key insight is that pseudo-labels are naturally imbalanced due to intrinsic data similarity, even when a model is trained on balanced source data and evaluated on balanced target data. If we address this previously unknown imbalanced classification problem arising from pseudo-labels instead of ground-truth training labels, we could remove model biases towards false majorities created by pseudo-labels. We propose a novel and effective debiased learning method with pseudo-labels, based on counterfactual reasoning and adaptive margins: The former removes the classifier response bias, whereas the latter adjusts the margin of each class according to the imbalance of pseudo-labels. Validated by extensive experimentation, our simple debiased learning delivers significant accuracy gains over the state-of-the-art on ImageNet-1K: 26% for semi-supervised learning with 0.2% annotations and 9% for zero-shot learning. Our code is available at: https://github.com/frank-xwang/debiased-pseudo-labeling. Xudong Wang 0007, Zhirong Wu, Long Lian, Stella X. Yu |
CVPR | 4 |
| 2022 | Unsupervised Selective Labeling for More Effective Semi-supervised Learning
Xudong Wang 0007, Long Lian, Stella X. Yu |
ECCV (30) | 3 |
| 2022 | Complex-valued Butterfly Transform for Efficient Hyperspectral Image ProcessingabstractHyper-spectral imaging (HSI) is a critical remote sensing modality that captures high-resolution spectral information in addition to high spatial resolution. Due to its ability to capture rich information about the material properties of the target, HSI has found applications such as agriculture, ecological monitoring, urban planning, and medicine. However, HSI images tend to have high spatial resolution and many channels, thus requiring models that can integrate information over a large context and identify spectral signatures. HSI classification datasets are also highly imbalanced, sparsely labeled, and significantly smaller than standard vision datasets like ImageNet or CIFAR, thus motivating smaller models with high sample efficiency. This work combines the strengths of multi-scale representations and data-driven feature learning. We generalize the butterfly transform to a learned complex-valued butterfly layer, allowing for parameter-efficient extraction of hierarchical complex-valued features from 1D signals. This allows us to create a lean yet highly accurate hyperspectral image classification model. Benchmarked on the Indian Pines, ROSIS-03 Pavia University, and Salinas datasets, our method demonstrates accuracy on par with SSDGL while using$7\mathbf{x}$fewer parameters. On the most imbalanced and sparsely labeled dataset, our method outperforms SSDGL. Utkarsh Singhal, Stella X. Yu |
IJCNN | 2 |
| 2022 | Transformer for 3D Point CloudsabstractDeep neural networks are widely used for understanding 3D point clouds. At each point convolution layer, features are computed from local neighbourhoods of 3D points and combined for subsequent processing in order to extract semantic information. Existing methods adopt the same individual point neighborhoods throughout the network layers, defined by the same metric on the fixed input point coordinates. This common practice is easy to implement but not necessarily optimal. Ideally, local neighborhoods should be different at different layers, as more latent information is extracted at deeper layers. We propose a novel end-to-end approach to learn different non-rigid transformations of the input point cloud so that optimal local neighborhoods can be adopted at each layer. We propose both linear (affine) and non-linear (projective and deformable) spatial transformers for 3D point clouds. With spatial transformers on the ShapeNet part segmentation dataset, the network achieves higher accuracy for all categories, with 8 percent gain on earphones and rockets in particular. Our method also outperforms the state-of-the-art on other point cloud tasks such as classification, detection, and semantic segmentation. Visualizations show that spatial transformers can learn features more efficiently by dynamically altering local neighborhoods according to the geometry and semantics of 3D shapes in spite of their within-category variations. Rudrasis Chakraborty, Stella X. Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | SurReal: Complex-Valued Learning as Principled Transformations on a Scaling and Rotation ManifoldabstractComplex-valued data are ubiquitous in signal and image processing applications, and complex-valued representations in deep learning have appealing theoretical properties. While these aspects have long been recognized, complex-valued deep learning continues to lag far behind its real-valued counterpart. We propose a principled geometric approach to complex-valued deep learning. Complex-valued data could often be subject to arbitrary complex-valued scaling; as a result, real and imaginary components could covary. Instead of treating complex values as two independent channels of real values, we recognize their underlying geometry: we model the space of complex numbers as a product manifold of nonzero scaling and planar rotations. Arbitrary complex-valued scaling naturally becomes a group of transitive actions on this manifold. We propose to extend the property instead of the form of real-valued functions to the complex domain. We define convolution as the weighted Fréchet mean on the manifold that is equivariant to the group of scaling/rotation actions and define distance transform on the manifold that is invariant to the action group. The manifold perspective also allows us to define nonlinear activation functions, such as tangent ReLU and G -transport, as well as residual connections on the manifold-valued data. We dub our model SurReal, as our experiments on MSTAR and RadioML deliver high performance with only a fractional size of real- and complex-valued baseline models. Rudrasis Chakraborty, Yifei Xing 0001, Stella X. Yu |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | Tied Block Convolution: Leaner and Better CNNs with Shared Thinner FiltersabstractConvolution is the main building block of a convolutional neural network (CNN). We observe that an optimized CNN often has highly correlated filters as the number of channels increases with depth, reducing the expressive power of feature representations. We propose Tied Block Convolution (TBC) that shares the same thinner filter over equal blocks of channels and produces multiple responses with a single filter. The concept of TBC can also be extended to group convolution and fully connected layers, and can be applied to various backbone networks and attention modules. Our extensive experimentation on classification, detection, instance segmentation, and attention demonstrates that TBC is consistently leaner and significantly better than standard convolution and group convolution. On attention, with 64 times fewer parameters, our TiedSE performs on par with the standard SE. On detection and segmentation, TBC can effectively handle highly overlapping instances, whereas standard CNNs often fail to accurately aggregate information in the presence of occlusion and result in multiple redundant partial object proposals. By sharing filters across channels, TBC reduces correlation and delivers a sizable gain of 6% in the average precision for object detection on MS-COCO when the occlusion ratio is 80%. Xudong Wang 0007, Stella X. Yu |
AAAI | 2 |
| 2021 | Unsupervised Feature Learning by Cross-Level Instance-Group DiscriminationabstractUnsupervised feature learning has made great strides with contrastive learning based on instance discrimination and invariant mapping, as benchmarked on curated class-balanced datasets. However, natural data could be highly correlated and long-tail distributed. Natural between-instance similarity conflicts with the presumed instance distinction, causing unstable training and poor performance.Our idea is to discover and integrate between-instance similarity into contrastive learning, not directly by instance grouping, but by cross-level discrimination (CLD) between instances and local instance groups. While invariant mapping of each instance is imposed by attraction within its augmented views, between-instance similarity could emerge from common repulsion against instance groups.Our batch-wise and cross-view comparisons also greatly improve the positive/negative sample ratio of contrastive learning and achieve better invariant mapping. To effect both grouping and discrimination objectives, we impose them on features separately derived from a shared representation. In addition, we propose normalized projection heads and unsupervised hyper-parameter tuning for the first time.Our extensive experimentation demonstrates that CLD is a lean and powerful add-on to existing methods such as NPID, MoCo, InfoMin, and BYOL on highly correlated, long-tail, or balanced datasets. It not only achieves new state-of-the-art on self-supervision, semi-supervision, and transfer learning benchmarks, but also beats MoCo v2 and SimCLR on every reported performance attained with a much larger compute. CLD effectively brings unsupervised learning closer to natural data and real-world applications. Our code is publicly available at: https://github.com/frank-xwang/CLD-UnsupervisedLearning. Xudong Wang 0007, Ziwei Liu 0002, Stella X. Yu |
CVPR | 3 |
| 2021 | Unsupervised Visual Attention and Invariance for Reinforcement LearningabstractVision-based reinforcement learning (RL) is successful, but how to generalize it to unknown test environments remains challenging. Existing methods focus on training an RL policy that is universal to changing visual domains, whereas we focus on extracting visual foreground that is universal, feeding clean invariant vision to the RL policy learner. Our method is completely unsupervised, without manual annotations or access to environment internals.Given videos of actions in a training environment, we learn how to extract foregrounds with unsupervised keypoint detection, followed by unsupervised visual attention to automatically generate a foreground mask per video frame. We can then introduce artificial distractors and train a model to reconstruct the clean foreground mask from noisy observations. Only this learned model is needed during test to provide distraction-free visual input to the RL policy learner.Our Visual Attention and Invariance (VAI) method significantly outperforms the state-of-the-art on visual domain generalization, gaining 15~49% (61~229%) more cumulative rewards per episode on DeepMind Control (our DrawerWorld Manipulation) benchmarks. Our results demonstrate that it is not only possible to learn domain-invariant vision without any supervision, but freeing RL from visual distractions also makes the policy more focused and thus far better. Xudong Wang 0007, Long Lian, Stella X. Yu |
CVPR | 3 |
| 2021 | Unsupervised Discriminative Learning of Sounds for Audio Event ClassificationabstractRecent progress in network-based audio event classification has shown the benefit of pre-training models on visual data such as ImageNet. While this process allows knowledge transfer across different domains, training a model on large-scale visual datasets is time consuming. On several audio event classification benchmarks, we show a fast and effective alternative that pre-trains the model unsupervised, only on audio data and yet delivers on-par performance with ImageNet pre-training. Furthermore, we show that our discriminative audio learning can be used to transfer knowledge across audio datasets and optionally include ImageNet pre-training. Sascha Hornauer, Kenneth Li 0002, Stella X. Yu, Shabnam Ghaffarzadegan, Liu Ren 0001 |
ICASSP | 3 |
| 2021 | Universal Weakly Supervised Segmentation by Pixel-to-Segment Contrastive Learning
Tsung-Wei Ke, Jyh-Jing Hwang, Stella X. Yu |
ICLR | 3 |
| 2021 | Long-tailed Recognition by Routing Diverse Distribution-Aware Experts
Xudong Wang 0007, Long Lian, Zhongqi Miao, Ziwei Liu 0002, Stella X. Yu |
ICLR | 5 |
| 2021 | Memory-Efficient Learning for High-Dimensional MRI Reconstruction
Ke Wang 0067, Michael R. Kellman, Christopher M. Sandino, Kevin Zhang 0003, S. S. Vasanawala, Jonathan I. Tamir, Stella X. Yu, Michael Lustig |
MICCAI (6) | 7 |
| 2021 | The Emergence of Objectness: Learning Zero-shot Segmentation from VideosabstractHumans can easily detect and segment moving objects simply by observing how they move, even without knowledge of object semantics. Inspired by this, we develop a zero-shot unsupervised approach for learning object segmentations. The model comprises two visual pathways: an appearance pathway that segments individual RGB images into coherent object regions, and a motion pathway that predicts the flow vector for each region between consecutive video frames. The two pathways jointly reconstruct a new representation called segment flow. This decoupled representation of appearance and motion is trained in a self-supervised manner to reconstruct one frame from another.When pretrained on an unlabeled video corpus, the model can be useful for a variety of applications, including 1) primary object segmentation from a single image in a zero-shot fashion; 2) moving object segmentation from a video with unsupervised test-time adaptation; 3) image semantic segmentation by supervised fine-tuning on a labeled image dataset. We demonstrate encouraging experimental results on all of these tasks using pretrained models. Runtao Liu, Zhirong Wu, Stella X. Yu, Stephen Lin 0001 |
NeurIPS | 3 |
| 2020 | Open Compound Domain AdaptationabstractA typical domain adaptation approach is to adapt models trained on the annotated data in a source domain (e.g., sunny weather) for achieving high performance on the test data in a target domain (e.g., rainy weather). Whether the target contains a single homogeneous domain or multiple heterogeneous domains, existing works always assume that there exist clear distinctions between the domains, which is often not true in practice (e.g., changes in weather). We study an open compound domain adaptation (OCDA) problem, in which the target is a compound of multiple homogeneous domains without domain labels, reflecting realistic data collection from mixed and novel situations. We propose a new approach based on two technical insights into OCDA: 1) a curriculum domain adaptation strategy to bootstrap generalization across domains in a data-driven self-organizing fashion and 2) a memory module to increase the model's agility towards novel domains. Our experiments on digit classification, facial expression recognition, semantic segmentation, and reinforcement learning demonstrate the effectiveness of our approach. Ziwei Liu 0002, Zhongqi Miao, Xingang Pan, Xiaohang Zhan, Dahua Lin, Stella X. Yu, Boqing Gong |
CVPR | 6 |
| 2020 | Orthogonal Convolutional Neural NetworksabstractDeep convolutional neural networks are hindered by training instability and feature redundancy towards further performance improvement. A promising solution is to impose orthogonality on convolutional filters. We develop an efficient approach to impose filter orthogonality on a convolutional layer based on the doubly block-Toeplitz matrix representation of the convolutional kernel, instead of the common kernel orthogonality approach, which we show is only necessary but not sufficient for ensuring orthogonal convolutions. Our proposed orthogonal convolution requires no additional parameters and little computational overhead. It consistently outperforms the kernel orthogonality alternative on a wide range of tasks such as image classification and inpainting under supervised, semi-supervised and unsupervised settings. It learns more diverse and expressive features with better training stability, robustness, and generalization. Our code is publicly available. Yubei Chen, Rudrasis Chakraborty, Stella X. Yu |
CVPR | 4 |
| 2020 | Unsupervised Sketch to Photo Synthesis
Runtao Liu, Stella X. Yu |
ECCV (3) | 3 |
| 2020 | BatVision: Learning to See 3D Spatial Layout with Two EarsabstractMany species have evolved advanced non-visual perception while artificial systems fall behind. Radar and ultrasound complement camera-based vision but they are often too costly and complex to set up for very limited information gain. In nature, sound is used effectively by bats, dolphins, whales, and humans for navigation and communication. However, it is unclear how to best harness sound for machine perception.Inspired by bats' echolocation mechanism, we design a low- cost BatVision system that is capable of seeing the 3D spatial layout of space ahead by just listening with two ears. Our system emits short chirps from a speaker and records returning echoes through microphones in an artificial human pinnae pair. During training, we additionally use a stereo camera to capture color images for calculating scene depths. We train a model to predict depth maps and even grayscale images from the sound alone. During testing, our trained BatVision provides surprisingly good predictions of 2D visual scenes from two 1D audio signals. Such a sound to vision system would benefit robot navigation and machine vision, especially in low-light or no-light conditions. Our code and data are publicly available. Jesper Haahr Christensen, Sascha Hornauer, Stella X. Yu |
ICRA | 3 |
| 2019 | Large-Scale Long-Tailed Recognition in an Open WorldabstractReal world data often have a long-tailed and open-ended distribution. A practical recognition system must classify among majority and minority classes, generalize from a few known instances, and acknowledge novelty upon a never seen instance. We define Open Long-Tailed Recognition (OLTR) as learning from such naturally distributed data and optimizing the classification accuracy over a balanced test set which include head, tail, and open classes. OLTR must handle imbalanced classification, few-shot learning, and open-set recognition in one integrated algorithm, whereas existing classification approaches focus only on one aspect and deliver poorly over the entire class spectrum. The key challenges are how to share visual knowledge between head and tail classes and how to reduce confusion between tail and open classes. We develop an integrated OLTR algorithm that maps an image to a feature space such that visual concepts can easily relate to each other based on a learned metric that respects the closed-world classification while acknowledging the novelty of the open world. Our so-called dynamic meta-embedding combines a direct image feature and an associated memory feature, with the feature norm indicating the familiarity to known classes. On three large-scale OLTR datasets we curate from object-centric ImageNet, scene-centric Places, and face-centric MS1M data, our method consistently outperforms the state-of-the-art. Our code, datasets, and models enable future OLTR research and are publicly available at \url{https://liuziwei7.github.io/projects/LongTail.html}. Ziwei Liu 0002, Zhongqi Miao, Xiaohang Zhan, Boqing Gong, Stella X. Yu |
CVPR | 6 |
| 2019 | Adversarial Structure Matching for Structured Prediction TasksabstractPixel-wise losses, i.e., cross-entropy or L2, have been widely used in structured prediction tasks as a spatial extension of generic image classification or regression. However, its i.i.d. assumption neglects the structural regularity present in natural images. Various attempts have been made to incorporate structural reasoning mostly through structure priors in a cooperative way where co-occurring patterns are encouraged. We, on the other hand, approach this problem from an opposing angle and propose a new framework, Adversarial Structure Matching (ASM), for training such structured prediction networks via an adversarial process, in which we train a structure analyzer that provides the supervisory signals, the ASM loss. The structure analyzer is trained to maximize ASM loss, or to emphasize recurring multi-scale hard negative structural mistakes usually among co-occurring patterns. On the contrary, the structured prediction network is trained to reduce those mistakes and is thus enabled to distinguish fine-grained structures. As a result, training structured prediction networks using ASM reduces contextual confusion among objects and improves boundary localization. We demonstrate that ASM outperforms its pixel-wise counterpart and commonly used structure priors, GAN, on three different structured prediction tasks, namely, semantic segmentation, monocular depth estimation, and surface normal prediction. Jyh-Jing Hwang, Tsung-Wei Ke, Jianbo Shi, Stella X. Yu |
CVPR | 4 |
| 2019 | SegSort: Segmentation by Discriminative Sorting of SegmentsabstractAlmost all existing deep learning approaches for semantic segmentation tackle this task as a pixel-wise classification problem. Yet humans understand a scene not in terms of pixels, but by decomposing it into perceptual groups and structures that are the basic building blocks of recognition. This motivates us to propose an end-to-end pixel-wise metric learning approach that mimics this process. In our approach, the optimal visual representation determines the right segmentation within individual images and associates segments with the same semantic classes across images. The core visual learning problem is therefore to maximize the similarity within segments and minimize the similarity between segments. Given a model trained this way, inference is performed consistently by extracting pixel-wise embeddings and clustering, with the semantic label determined by the majority vote of its nearest neighbors from an annotated set. As a result, we present the SegSort, as a first attempt using deep learning for unsupervised semantic segmentation, achieving 76% performance of its supervised counterpart. When supervision is available, SegSort shows consistent improvements over conventional approaches based on pixel-wise softmax training. Additionally, our approach produces more precise boundaries and consistent region predictions. The proposed SegSort further produces an interpretable result, as each choice of label can be easily understood from the retrieved nearest segments. Jyh-Jing Hwang, Stella X. Yu, Jianbo Shi, Maxwell D. Collins, Tien-Ju Yang, Liang-Chieh Chen |
ICCV | 2 |
| 2018 | Unsupervised Feature Learning via Non-Parametric Instance DiscriminationabstractNeural net classifiers trained on data with annotated class labels can also capture apparent visual similarity among categories without being directed to do so. We study whether this observation can be extended beyond the conventional domain of supervised learning: Can we learn a good feature representation that captures apparent similarity among instances, instead of classes, by merely asking the feature to be discriminative of individual instances? We formulate this intuition as a non-parametric classification problem at the instance-level, and use noise-contrastive estimation to tackle the computational challenges imposed by the large number of instance classes. Our experimental results demonstrate that, under unsupervised learning settings, our method surpasses the state-of-the-art on ImageNet classification by a large margin. Our method is also remarkable for consistently improving test performance with more training data and better network architectures. By fine-tuning the learned feature, we further obtain competitive results for semi-supervised learning and object detection tasks. Our non-parametric model is highly compact: With 128 features per image, our method requires only 600MB storage for a million images, enabling fast nearest neighbour retrieval at the run time. Zhirong Wu, Yuanjun Xiong, Stella X. Yu, Dahua Lin |
CVPR | 3 |
| 2018 | Adaptive Affinity Fields for Semantic Segmentation
Tsung-Wei Ke, Jyh-Jing Hwang, Ziwei Liu 0002, Stella X. Yu |
ECCV (1) | 4 |
| 2018 | Improving Generalization via Scalable Neighborhood Component Analysis
Zhirong Wu, Alexei A. Efros, Stella X. Yu |
ECCV (7) | 3 |
| 2017 | Convolutional Random Walk Networks for Semantic Image SegmentationabstractMost current semantic segmentation methods rely on fully convolutional networks (FCNs). However, their use of large receptive fields and many pooling layers cause low spatial resolution inside the deep layers. This leads to predictions with poor localization around the boundaries. Prior work has attempted to address this issue by post-processing predictions with CRFs or MRFs. But such models often fail to capture semantic relationships between objects, which causes spatially disjoint predictions. To overcome these problems, recent methods integrated CRFs or MRFs into an FCN framework. The downside of these new models is that they have much higher complexity than traditional FCNs, which renders training and testing more challenging. In this work we introduce a simple, yet effective Convolutional Random Walk Network (RWN) that addresses the issues of poor boundary localization and spatially fragmented predictions with very little increase in model complexity. Our proposed RWN jointly optimizes the objectives of pixelwise affinity and semantic segmentation. It combines these two objectives via a novel random walk layer that enforces consistent spatial grouping in the deep layers of the network. Our RWN is implemented using standard convolution and matrix multiplication. This allows an easy integration into existing FCN frameworks and it enables end-to-end training of the whole network via standard back-propagation. Our implementation of RWN requires just 131 additional parameters compared to the traditional FCNs, and yet it consistently produces an improvement over the FCNs on semantic segmentation and scene labeling. Gedas Bertasius, Lorenzo Torresani, Stella X. Yu, Jianbo Shi |
CVPR | 3 |
| 2017 | Multigrid Neural Architectures
Tsung-Wei Ke, Michael Maire, Stella X. Yu |
CVPR | 3 |
| 2017 | Learning Non-Lambertian Object Intrinsics Across ShapeNet CategoriesabstractWe focus on the non-Lambertian object-level intrinsic problem of recovering diffuse albedo, shading, and specular highlights from a single image of an object. Based on existing 3D models in the ShapeNet database, a large-scale object intrinsics database is rendered with HDR environment maps. Millions of synthetic images of objects and their corresponding albedo, shading, and specular ground-truth images are used to train an encoder-decoder CNN, which can decompose an image into the product of albedo and shading components along with an additive specular component. Our CNN delivers accurate and sharp results in this classical inverse problem of computer vision. Evaluated on our realistically synthetic dataset, our method consistently outperforms the state-of-the-art by a large margin. We train and test our CNN across different object categories. Perhaps surprising especially from the CNN classification perspective, our intrinsics CNN generalizes very well across categories. Our analysis shows that feature learning at the encoder stage is more crucial for developing a universal representation across categories. We apply our model to real images and videos from Internet, and observe robust and realistic intrinsics results. Quality non-Lambertian intrinsics could open up many interesting applications such as realistic product search based on material properties and image-based albedo/specular editing. Yue Dong 0001, Stella X. Yu |
CVPR | 4 |
| 2017 | Unsupervised Learning of Important Objects from First-Person VideosabstractA first-person camera, placed at a person's head, captures, which objects are important to the camera wearer. Most prior methods for this task learn to detect such important objects from the manually labeled first-person data in a supervised fashion. However, important objects are strongly related to the camera wearer's internal state such as his intentions and attention, and thus, only the person wearing the camera can provide the importance labels. Such a constraint makes the annotation process costly and limited in scalability. In this work, we show that we can detect important objects in first-person images without the supervision by the camera wearer or even third-person labelers. We formulate an important detection problem as an interplay between the 1) segmentation and 2) recognition agents. The segmentation agent first proposes a possible important object segmentation mask for each image, and then feeds it to the recognition agent, which learns to predict an important object mask using visual semantics and spatial features. We implement such an interplay between both agents via an alternating cross-pathway supervision scheme inside our proposed Visual-Spatial Network (VSN). Our VSN consists of spatial ("where") and visual ("what") pathways, one of which learns common visual semantics while the other focuses on the spatial location cues. Our unsupervised learning is accomplished via a cross-pathway supervision, where one pathway feeds its predictions to a segmentation agent, which proposes a candidate important object segmentation mask that is then used by the other pathway as a supervisory signal. We show our method's success on two different important object datasets, where our method achieves similar or better results as the supervised methods. Gedas Bertasius, Hyun Soo Park, Stella X. Yu, Jianbo Shi |
ICCV | 3 |
| 2017 | Am I a Baller? Basketball Performance Assessment from First-Person VideosabstractThis paper presents a method to assess a basketball player's performance from his/her first-person video. A key challenge lies in the fact that the evaluation metric is highly subjective and specific to a particular evaluator. We leverage the first-person camera to address this challenge. The spatiotemporal visual semantics provided by a first-person view allows us to reason about the camera wearer's actions while he/she is participating in an unscripted basketball game. Our method takes a player's first-person video and provides a player's performance measure that is specific to an evaluator's preference. To achieve this goal, we first use a convolutional LSTM network to detect atomic basketball events from first-person videos. Our network's ability to zoom-in to the salient regions addresses the issue of a severe camera wearer's head movement in first-person videos. The detected atomic events are then passed through the Gaussian mixtures to construct a highly non-linear visual spatiotemporal basketball assessment feature. Finally, we use this feature to learn a basketball assessment model from pairs of labeled first-person basketball videos, for which a basketball expert indicates, which of the two players is better. We demonstrate that despite not knowing the basketball evaluator's criterion, our model learns to accurately assess the players in real-world games. Furthermore, our model can also discover basketball events that contribute positively and negatively to a player's performance. Gedas Bertasius, Hyun Soo Park, Stella X. Yu, Jianbo Shi |
ICCV | 3 |
| 2017 | Mooney face classification and prediction by learning across toneabstractMooney faces are special two-tone images that elicit a rich impression of identity and facial expression in human observers. While Mooney faces are important, there exist only a small number of instances hand-crafted from source photos which are often no longer available. We first apply deep learning methods to generate a plausible Mooney face automatically from any face photo. We are then able to create a large-scale face dataset with paired grayscale and two-tone images. We then study how well two-tone versions make face predictions, using conditional Generative Adversarial Networks. We show that faces predicted from Mooney images bear striking resemblance to source photos, and they are better than two-tone images obtained by global intensity thresholding. We also demonstrate remarkable face predictions from very low resolution surveillance photos. Our findings reveal great potentials of combining deep learning and Mooney faces for more effective face recognition in a wide range of conditions. Tsung-Wei Ke, Stella X. Yu, David Whitney |
ICIP | 2 |
| 2017 | Ground2sky label transfer for fine-grained aerial car recognitionabstractOverhead images captured by helicopters, unmanned aerial vehicles and satellites are widely available. Prior aerial target recognition methods mainly deal with generic object categories such as cars, roads, and boats. We go beyond this and aim for fine-grained recognition, e.g., distinguishing between a Toyota and a Honda sedan. This task is so challenging for human annotators that labeling images directly is no longer an option: annotators are often unable to identify the object from such an extreme viewpoint and at such a low resolution. We propose a novel solution to collect fine-grained annotations of aerial images and develop the first ground-to-sky cross-view car dataset with instance-level correspondences. We compare the performance of human experts and deep learning approaches on fine-grained car recognition from aerial imagery. Noting that intraclass variation in aerial images is limited, we further show that with simple data augmentation, a classifier can be trained from fewer instances yet achieves comparable or even significantly better performance than human experts. Our experimental evidence demonstrates that fine-grained object recognition from overhead images is not only feasible but also well suited for deep learning methods. Our dataset is available at: http://ai.bu.edu/Ground2Sky/. Baochen Sun, Xingchao Peng, Stella X. Yu, Kate Saenko |
ICIP | 3 |
| 2017 | Better than real: Complex-valued neural nets for MRI fingerprintingabstractThe task of MRI fingerprinting is to identify tissue parameters from complex-valued MRI signals. The prevalent approach is dictionary based, where a test MRI signal is compared to stored MRI signals with known tissue parameters and the most similar signals and tissue parameters retrieved. Such an approach does not scale with the number of parameters and is rather slow when the tissue parameter space is large. Our first novel contribution is to use deep learning as an efficient nonlinear inverse mapping approach. We generate synthetic (tissue, MRI) data from an MRI simulator, and use them to train a deep net to map the MRI signal to the tissue parameters directly. Our second novel contribution is to develop a complex-valued neural network with new cardioid activation functions. Our results demonstrate that complex-valued neural nets could be much more accurate than real-valued neural nets at complex-valued MRI fingerprinting. Patrick Virtue, Stella X. Yu, Michael Lustig |
ICIP | 2 |
| 2016 | The perception of symmetry in the moving image: multi-level computational analysis of cinematographic scene structure and its visual receptionabstractThis research is driven by visuo-spatial perception focussed cognitive film studies, where the key emphasis is on the systematic study and generation of evidence that can characterise and establish correlates between principles for the synthesis of the moving image, and its cognitive (e.g., embodied visuo-auditory, emotional) recipient effects on observers [Suchan and Bhatt 2016b; Suchan and Bhatt 2016a]. Within this context, we focus on the case of "symmetry" in the cinematographic structure of the moving image, and propose a multi-level model of interpreting symmetric patterns therefrom. This provides the foundation for integrating scene analysis with the analysis of its visuo-spatial perception based on eye-tracking data. This is achieved by the integration of: computational semantic interpretation of the scene [Suchan and Bhatt 2016b] ---involving scene objects (people, objects in the scene), cinematographic aids (camera movement, shot types, cuts and scene structure)--- and perceptual artefacts (fixations, saccades, scan-path, areas of attention). Jakob Suchan, Mehul Bhatt, Stella X. Yu |
SAP | 3 |
| 2016 | Affinity CNN: Learning Pixel-Centric Pairwise Relations for Figure/Ground EmbeddingabstractSpectral embedding provides a framework for solving perceptual organization problems, including image segmentation and figure/ground organization. From an affinity matrix describing pairwise relationships between pixels, it clusters pixels into regions, and, using a complex-valued extension, orders pixels according to layer. We train a convolutional neural network (CNN) to directly predict the pair-wise relationships that define this affinity matrix. Spectral embedding then resolves these predictions into a globally-consistent segmentation and figure/ground organization of the scene. Experiments demonstrate significant benefit to this direct coupling compared to prior works which use explicit intermediate stages, such as edge detection, on the pathway from image to affinities. Our results suggest spectral embedding as a powerful alternative to the conditional random field (CRF)-based globalization schemes typically coupled to deep neural networks. Michael Maire, Takuya Narihira, Stella X. Yu |
CVPR | 3 |
| 2016 | Fine-to-coarse knowledge transfer for low-res image classificationabstractWe address the difficult problem of distinguishing fine-grained object categories in low resolution images. We propose a simple an effective deep learning approach that transfers fine-grained knowledge gained from high resolution training data to the coarse low-resolution test scenario. Such fine-to-coarse knowledge transfer has many real world applications, such as identifying objects in surveillance photos or satellite images where the image resolution at the test time is very low but plenty of high resolution photos of similar objects are available. Our extensive experiments on two standard benchmark datasets containing fine-grained car models and bird species demonstrate that our approach can effectively transfer fine-detail knowledge to coarse-detail imagery. Xingchao Peng, Judy Hoffman, Stella X. Yu, Kate Saenko |
ICIP | 3 |
| 2015 | Learning lightness from human judgement on relative reflectanceabstractWe develop a new approach to inferring lightness, the perceived reflectance of surfaces, from a single image. Classic methods view this problem from the perspective of intrinsic image decomposition, where an image is separated into reflectance and shading components. Rather than reason about reflectance and shading together, we learn to directly predict lightness differences between pixels. Large-scale training from human judgement data on relative reflectance, and patch representations built using deep networks, provide the foundation for our model. Benchmarked on the Intrinsic Images in the Wild dataset [4], our local lightness model achieves on-par performance with the state-of-the-art global lightness model, which incorporates multiple shading/reflectance priors and simultaneous reasoning between pairs of pixels in a dense conditional random field formulation. Takuya Narihira, Michael Maire, Stella X. Yu |
CVPR | 3 |
| 2015 | FlowWeb: Joint image set alignment by weaving consistent, pixel-wise correspondencesabstractGiven a set of poorly aligned images of the same visual concept without any annotations, we propose an algorithm to jointly bring them into pixel-wise correspondence by estimating a FlowWeb representation of the image set. FlowWeb is a fully-connected correspondence flow graph with each node representing an image, and each edge representing the correspondence flow field between a pair of images, i.e. a vector field indicating how each pixel in one image can find a corresponding pixel in the other image. Correspondence flow is related to optical flow but allows for correspondences between visually dissimilar regions if there is evidence they correspond transitively on the graph. Our algorithm starts by initializing all edges of this complete graph with an off-the-shelf, pairwise flow method. We then iteratively update the graph to force it to be more self-consistent. Once the algorithm converges, dense, globally-consistent correspondences can be read off the graph. Our results suggest that FlowWeb improves alignment accuracy over previous pairwise as well as joint alignment methods. Tinghui Zhou, Yong Jae Lee, Stella X. Yu, Alexei A. Efros |
CVPR | 3 |
| 2015 | Direct Intrinsics: Learning Albedo-Shading Decomposition by Convolutional RegressionabstractWe introduce a new approach to intrinsic image decomposition, the task of decomposing a single image into albedo and shading components. Our strategy, which we term direct intrinsics, is to learn a convolutional neural network (CNN) that directly predicts output albedo and shading channels from an input RGB image patch. Direct intrinsics is a departure from classical techniques for intrinsic image decomposition, which typically rely on physically-motivated priors and graph-based inference algorithms. The large-scale synthetic ground-truth of the MPI Sintel dataset plays the key role in training direct intrinsics. We demonstrate results on both the synthetic images of Sintel and the real images of the classic MIT intrinsic image dataset. On Sintel, direct intrinsics, using only RGB input, outperforms all prior work, including methods that rely on RGB+Depth input. Direct intrinsics also generalizes across modalities, our Sintel-trained CNN produces quite reasonable decompositions on the real images of the MIT dataset. Our results indicate that the marriage of CNNs with synthetic training data may be a powerful new technique for tackling classic problems in computer vision. Takuya Narihira, Michael Maire, Stella X. Yu |
ICCV | 3 |
| 2014 | Reconstructive Sparse Code Transfer for Contour Detection and Semantic Labeling
Michael Maire, Stella X. Yu, Pietro Perona |
ACCV (4) | 2 |
| 2013 | Hierarchical Scene AnnotationabstractWe present a computer-assisted annotation system, together with a labeled dataset and benchmark suite, for evaluating an algorithm's ability to recover hierarchical scene structure.We evolve segmentation groundtruth from the two-dimensional image partition into a tree model that captures both occlusion and object-part relationships among possibly overlapping regions.Our tree model extends the segmentation problem to encompass object detection, object-part containment, and figure-ground ordering.We mitigate the cost of providing richer groundtruth labeling through a new webbased annotation tool with an intuitive graphical interface for rearranging the region hierarchy.Using precomputed superpixels, our tool also guides creation of user-specified regions with pixel-perfect boundaries.Widespread adoption of this human-machine combination should make the inaccuracies of bounding box labeling a relic of the past.Evaluating the state-of-the-art in fully automatic image segmentation reveals that it produces accurate two-dimension partitions, but does not respect groundtruth object-part structure.Our dataset and benchmark is the first to quantify these inadequacies.We illuminate recovery of rich scene structure as an important new goal for segmentation. Michael Maire, Stella X. Yu, Pietro Perona |
BMVC | 2 |
| 2013 | Progressive Multigrid Eigensolvers for Multiscale Spectral SegmentationabstractWe reexamine the role of multiscale cues in image segmentation using an architecture that constructs a globally coherent scale-space output representation. This characteristic is in contrast to many existing works on bottom-up segmentation, which prematurely compress information into a single scale. The architecture is a standard extension of Normalized Cuts from an image plane to an image pyramid, with cross-scale constraints enforcing consistency in the solution while allowing emergence of coarse-to-fine detail. We observe that multiscale processing, in addition to improving segmentation quality, offers a route by which to speed computation. We make a significant algorithmic advance in the form of a custom multigrid eigensolver for constrained Angular Embedding problems possessing coarse-to-fine structure. Multiscale Normalized Cuts is a special case. Our solver builds atop recent results on randomized matrix approximation, using a novel interpolation operation to mold its computational strategy according to cross-scale constraints in the problem definition. Applying our solver to multiscale segmentation problems demonstrates speedup by more than an order of magnitude. This speedup is at the algorithmic level and carries over to any implementation target. Michael Maire, Stella X. Yu |
ICCV | 2 |
| 2012 | Power SVM: Generalization with exemplar classification uncertaintyabstractThe human vision tends to recognize more variants of a distinctive exemplar. This observation suggests that discriminative power of training exemplars could be utilized for shaping a desirable global classifier that generalizes maximally from a few exemplars. We propose to derive classification uncertainty for each exemplar, using a local classification task to separate the exemplar from those in other categories. We then design a global classifier by incorporating these uncertainties into constraints on the classifier margins. We show through the dual form that the classification criterion can be interpreted as finding closest points between convex hulls in the feature space augmented by classification uncertainty. We call this scheme Power SVM (as in Power Diagram), since each exemplar is no longer a singular point in the feature space, but a super-point with its own governing power in the classifier space. We test Power SVM on digit recognition, indoor-outdoor categorization, and large-scale scene classification tasks. It shows consistent improvement over SVM and uncertainty weighted SVM, especially when the number of training exemplars is small. Stella X. Yu, Shang-Hua Teng |
CVPR | 2 |
| 2012 | Angular Embedding: A Robust Quadratic CriterionabstractGiven the size and confidence of pairwise local orderings, angular embedding (AE) finds a global ordering with a near-global optimal eigensolution. As a quadratic criterion in the complex domain, AE is remarkably robust to outliers, unlike its real domain counterpart LS, the least squares embedding. Our comparative study of LS and AE reveals that AE's robustness is due not to the particular choice of the criterion, but to the choice of representation in the complex domain. When the embedding is encoded in the angular space, we not only have a nonconvex error function that delivers robustness, but also have a Hermitian graph Laplacian that completely determines the optimum and delivers efficiency. The high quality of embedding by AE in the presence of outliers can hardly be matched by LS, its corresponding L(1) norm formulation, or their bounded versions. These results suggest that the key to overcoming outliers lies not with additionally imposing constraints on the embedding solution, but with adaptively penalizing inconsistency between measurements themselves. AE thus significantly advances statistical ranking methods by removing the impact of outliers directly without explicit inconsistency characterization, and advances spectral clustering methods by covering the entire size-confidence measurement space and providing an ordered cluster organization. Stella X. Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2011 | Object detection and segmentation from joint embedding of parts and pixelsabstractWe present a new framework in which image segmentation, figure/ground organization, and object detection all appear as the result of solving a single grouping problem. This framework serves as a perceptual organization stage that integrates information from low-level image cues with that of high-level part detectors. Pixels and parts each appear as nodes in a graph whose edges encode both affinity and ordering relationships. We derive a generalized eigen-problem from this graph and read off an interpretation of the image from the solution eigenvectors. Combining an off-the-shelf top-down part-based person detector with our low-level cues and grouping formulation, we demonstrate improvements to object detection and segmentation. Michael Maire, Stella X. Yu, Pietro Perona |
ICCV | 2 |
| 2011 | Pop out many small structures from a very large microscopic image
Elena Bernardis, Stella X. Yu |
Medical Image Anal. | 2 |
| 2011 | Linear Scale and Rotation Invariant MatchingabstractMatching visual patterns that appear scaled, rotated, and deformed with respect to each other is a challenging problem. We propose a linear formulation that simultaneously matches feature points and estimates global geometrical transformation in a constrained linear space. The linear scheme enables search space reduction based on the lower convex hull property so that the problem size is largely decoupled from the original hard combinatorial problem. Our method therefore can be used to solve large scale problems that involve a very large number of candidate feature points. Without using prepruning in the search, this method is more robust in dealing with weak features and clutter. We apply the proposed method to action detection and image matching. Our results on a variety of images and videos demonstrate that our method is accurate, efficient, and robust. Hao Jiang 0007, Stella X. Yu, David R. Martin 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2010 | Finding dots: Segmentation as popping out regions from boundariesabstractMany applications need to segment out all small round regions in an image. This task of finding dots can be viewed as a region segmentation problem where the dots form one region and the areas between dots form the other. We formulate it as a graph cuts problem with two types of grouping cues: short-range attraction based on feature similarity and long-range repulsion based on feature dissimilarity. The feature we use is a pixel-centric relational representation that encodes local convexity: Pixels inside the dots and outside the dots become sinks and sources of the feature vector. Normalized cuts on both attraction and repulsion pop out all the dots in a single binary segmentation. Our experiments show that our method is more accurate and robust than state-of-art segmentation algorithms on four categories of microscopic images. It can also detect textons in natural scene images with the same set of parameters. Elena Bernardis, Stella X. Yu |
CVPR | 2 |
| 2010 | Classification and feature selection with human performance dataabstractWe investigate the utility of a novel form of prior, namely the accuracies with which humans categorize briefly displayed images. Such information reflects the complexity of an image for the visual system and carries information about the features important for categorization. We incorporate the prior in an SVM framework, by biasing the decision boundary towards examples difficult for humans, and by learning a suitable kernel. We focus on the task indoors vs. outdoors using a variety of histogram and interest point features. We observe improvement in classification especially for the indoor class when gist features are used. Christina Pavlopoulou, Stella X. Yu |
ICIP | 2 |
| 2010 | Segmentation Subject to Stitching Constraints: Finding Many Small Structures in a Large Image
Elena Bernardis, Stella X. Yu |
MICCAI (1) | 2 |
| 2010 | Feature Transitions with Saccadic Search: Size, Color, and Orientation Are Not AlikeabstractSize, color, and orientation have long been considered elementary features whose attributes are extracted in parallel and available to guide the deployment of attention. If each is processed in the same fashion with simply a different set of local detectors, one would expect similar search behaviours on localizing an equivalent flickering change among identically laid out disks. We analyze feature transitions associated with saccadic search and find out that size, color, and orientation are not alike in dynamic attribute processing over time. The Markovian feature transition is attractive for size, repulsive for color, and largely reversible for orientation. Stella X. Yu |
NIPS | 1 |
| 2009 | Linear solution to scale and rotation invariant object matchingabstractImages of an object undergoing ego- or camera-motion often appear to be scaled, rotated, and deformed versions of each other. To detect and match such distorted patterns to a single sample view of the object requires solving a hard computational problem that has eluded most object matching methods. We propose a linear formulation that simultaneously finds feature point correspondences and global geometrical transformations in a constrained solution space. Further reducing the search space based on the lower convex hull property of the formulation, our method scales well with the number of candidate features. Our results on a variety of images and videos demonstrate that our method is accurate, efficient, and robust over local deformation, occlusion, clutter, and large geometrical transformations. Hao Jiang 0007, Stella X. Yu |
CVPR | 2 |
| 2009 | Angular embedding: From jarring intensity differences to perceived luminanceabstractOur goal is to turn an intensity image into its perceived luminance without parsing it into depths, surfaces, or scene illuminations. We start with jarring intensity differences at two scales mixed according to edges, identified by a pixel-centric edge detector. We propose angular embedding as a more robust, efficient, and versatile alternative to LS, LLE, and NCUTS for obtaining a global brightness ordering from local differences. Our model explains a variety of brightness illusions with a single algorithm. Brightness of a pixel can be understood locally as its intensity deviating in the gradient direction and globally as finding its rank relative to others, particularly the lightest and darkest ones. Stella X. Yu |
CVPR | 1 |
| 2005 | Segmentation Induced by Scale InvarianceabstractPerceptual organization is scale-invariant. In turn, a segmentation that separates features consistently at all scales is the desired one that reveals the underlying structural organization of an image. Addressing cross-scale correspondence with interior pixels, we develop this intuition into a general segmenter that handles texture and illusory contours through edges entirely without any explicit characterization of texture or curvilinearity. Experimental results demonstrate that our method not only performs on par with either texture segmentation or boundary completion methods on their specialized examples, but also works well on a variety of real images. Stella X. Yu |
CVPR (1) | 1 |
| 2004 | Segmentation Using Multiscale Cues
Stella X. Yu |
CVPR (1) | 1 |
| 2004 | Segmentation Given Partial Grouping ConstraintsabstractWe consider data clustering problems where partial grouping is known a priori. We formulate such biased grouping problems as a constrained optimization problem, where structural properties of the data define the goodness of a grouping and partial grouping cues define the feasibility of a grouping. We enforce grouping smoothness and fairness on labeled data points so that sparse partial grouping information can be effectively propagated to the unlabeled data. Considering the normalized cuts criterion in particular, our formulation leads to a constrained eigenvalue problem. By generalizing the Rayleigh-Ritz theorem to projected matrices, we find the global optimum in the relaxed continuous domain by eigendecomposition, from which a near-global optimum to the discrete labeling problem can be obtained effectively. We apply our method to real image segmentation problems, where partial grouping priors can often be derived based on a crude spatial attentional map that binds places with common salient features or focuses on expected object locations. We demonstrate not only that it is possible to integrate both image structures and priors in a single grouping process, but also that objects can be segregated from the background without specific object knowledge. Stella X. Yu, Jianbo Shi |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2003 | Object-Specific Figure-Ground Segregation
Stella X. Yu, Jianbo Shi |
CVPR (2) | 1 |
| 2003 | Multiclass Spectral ClusteringabstractWe propose a principled account on multiclass spectral clustering. Given a discrete clustering formulation, we first solve a relaxed continuous optimization problem by eigen-decomposition. We clarify the role of eigenvectors as a generator of all optimal solutions through orthonormal transforms. We then solve an optimal discretization problem, which seeks a discrete solution closest to the continuous optima. The discretization is efficiently computed in an iterative fashion using singular value decomposition and nonmaximum suppression. The resulting discrete solutions are nearly global-optimal. Our method is robust to random initialization and converges faster than other clustering methods. Experiments on real image segmentation are reported. Stella X. Yu, Jianbo Shi |
ICCV | 1 |
| 2002 | Concurrent Object Recognition and Segmentation by Graph PartitioningabstractSegmentation and recognition have long been treated as two separate pro(cid:173) cesses. We propose a mechanism based on spectral graph partitioning that readily combine the two processes into one. A part-based recogni(cid:173) tion system detects object patches, supplies their partial segmentations as well as knowledge about the spatial configurations of the object. The goal of patch grouping is to find a set of patches that conform best to the object configuration, while the goal of pixel grouping is to find a set of pixels that have the best low-level feature similarity. Through pixel-patch in(cid:173) teractions and between-patch competition encoded in the solution space, these two processes are realized in one joint optimization problem. The globally optimal partition is obtained by solving a constrained eigenvalue problem. We demonstrate that the resulting object segmentation elimi(cid:173) nates false positives for the part detection, while overcoming occlusion and weak contours for the low-level edge detection. Stella X. Yu, Ralph Gross, Jianbo Shi |
NIPS | 1 |
| 2001 | Understanding Popout through RepulsionabstractPerceptual popout is defined by both feature similarity and local feature contrast. We identify these two measures with attraction and repulsion, and unify the dual processes of association by attraction and segregation by repulsion in a single grouping framework. We generalize normalized cuts to multi-way partitioning with these dual measures. We expand graph partitioning approaches to weight matrices with negative entries, and provide a theoretical basis for solution regularization in such algorithms. We show that attraction, repulsion and regularization each contributes in a unique way to popout. Their roles are demonstrated in various salience detection and visual search scenarios. This work opens up the possibilities of encoding negative correlations in constraint satisfaction problems, where solutions by simple and robust eigendecomposition become possible. Stella X. Yu, Jianbo Shi |
CVPR (2) | 1 |
| 2001 | Segmentation with Pairwise Attraction and Repulsion
Stella X. Yu, Jianbo Shi |
ICCV | 1 |
| 2001 | Grouping with BiasabstractWith the optimization of pattern discrimination as a goal, graph partitioning approaches often lack the capability to integrate prior knowledge to guide grouping. In this paper, we consider priors from unitary generative models, partially labeled data and spatial attention. These priors are modelled as constraints in the solution space. By imposing uniformity condition on the constraints, we restrict the feasible space to one of smooth solutions. A subspace projection method is developed to solve this constrained eigenprob(cid:173) lema We demonstrate that simple priors can greatly improve image segmentation results. Stella X. Yu, Jianbo Shi |
NIPS | 1 |
| 2000 | What do V1 neurons tell us about saccadic suppression?
Stella X. Yu, Tai Sing Lee |
Neurocomputing | 1 |
| 1999 | An Information-Theoretic Framework for Understanding Saccadic Eye Movements
Tai Sing Lee, Stella X. Yu |
NIPS | 2 |