Yin Li 0003

dblp:49/5981-3 · DBLP profile ↗
← Back
84ranked-venue papers
16as first author
39since 2021 · last 2026
0000-0003-4173-9453ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 63 · 13 first-author · 26 since 2021Artificial intelligence and machine learning · 62 · 11 first-author · 32 since 2021Computer networks · 3 · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 ApproxBit: Efficient Video Analytics through Latency-Aware Offloading with Learned Binary Codes
abstract
With the growing ubiquity of video content, efficient video analytics has become essential for applications such as surveillance, autonomous driving, and augmented reality. Yet, deploying video analytics models on resource-constrained edge devices and in low-bandwidth environments remains challenging. A dominant method for handling demanding video analytics tasks on edge devices has been to offload computation strategically from the edge device to servers. However, all prior solutions fail to offload under severely constrained, real-world network conditions (such as, a few-Mbps satellite network) due to the much higher data rates associated with video tasks. We introduce ApproxBit, a system to optimize shared edge-to-cloud processing for video analytics tasks; the two that we experiment with are video action recognition and video question answering. ApproxBit integrates an encoder within the video model, uses learned binary codes to effectively compress and offload data, and adaptively decides on the offloading point depending on the network bandwidth. ApproxBit’s adaptive and efficient data compression, which reduces the original feature map size by up to 2142.4 ×, makes it an ideal solution for video analytics on edge devices, especially with constrained networks. We evaluate ApproxBit on the two video tasks, across different model architectures (e.g., convolution- and Transformer-based) and multiple datasets (e.g., Something-Something-v2, Kinetics, and MSVD). Our results of latency and accuracy are superior over baselines: edge-only processing, server-only processing, DNN Surgery [ToCC ’23], full offloading of H.264-encoded videos, DeepCOD [SenSys ’20], neural video compression DCVC-FM [CVPR ’24], and LimitNet [MobiSys ’24]. We also demonstrate ApproxBit’s adaptivity to changing network conditions, and generalization in a real-world user study.
Hyunseung Kim, Sheetal Prasanna, Yin Li 0003, Somali Chaterji, Saurabh Bagchi
SenSys3
2025 PAVE: Patching and Adapting Video Large Language Models
abstract
Pre-trained video large language models (Video LLMs) exhibit remarkable reasoning capabilities, yet adapting these models to new tasks involving additional modalities or data types (e.g., audio or 3D information) remains challenging. In this paper, we present PAVE, a flexible framework for adapting pre-trained Video LLMs to downstream tasks with side-channel signals, such as audio, 3D cues, or multi-view videos. PAVE introduces lightweight adapters, referred to as "patches," which adds a small number of parameters and operations to a base model without modifying its architecture or pre-trained weights. In doing so, PAVE can effectively adapt the pre-trained base model to support diverse downstream tasks, including audio-visual question answering, 3D reasoning, multi-view video recognition, and high frame rate video understanding. Across these tasks, PAVE significant enhances the performance of the base model, surpassing state-of-the-art task-specific models while incurring a minor cost of ~0.1% additional FLOPs and parameters. Further, PAVE supports multitask learning and generalizes well across different Video LLMs. Our code is available at https://github.com/dragonlzm/PAVE.
Zhuoming Liu 0001, Yiquan Li, Khoi D. Nguyen 0001, Yiwu Zhong, Yin Li 0003
CVPR5
2025 Robust 3D Object Detection Using Probabilistic Point Clouds From Single-Photon Lidars
abstract
LiDAR-based 3D sensors provide point clouds, a canonical 3D representation used in various scene understanding tasks. Modern LiDARs face key challenges in several real-world scenarios, such as long-distance or low-albedo objects, producing sparse or erroneous point clouds. These errors, which are rooted in the noisy raw LiDAR measurements, get propagated to downstream perception models, resulting in potentially severe loss of accuracy. This is because conventional 3D processing pipelines do not retain any uncertainty information from the raw measurements when constructing point clouds. We propose Probabilistic Point Clouds (PPC), a novel 3D scene representation where each point is augmented with a probability attribute that encapsulates the measurement uncertainty (or confidence) in the raw data. We further introduce inference approaches that leverage PPC for robust 3D object detection; these methods are versatile and can be used as computationally lightweight drop-in modules in 3D inference pipelines. We demonstrate, via both simulations and real captures, that PPC-based 3D inference methods outperform several baselines using LiDAR as well as camera-LiDAR fusion models, across challenging indoor and outdoor scenarios involving small, distant, and low-albedo objects, as well as strong ambient light. Our project webpage is at https://bhavyagoyal.github.io/ppc .
Bhavya Goyal, Felipe Gutierrez-Barragan, Andreas Velten, Yin Li 0003, Mohit Gupta 0001
ICCV5
2025 Recovering Parametric Scenes from Very Few Time-of-Flight Pixels
abstract
We aim to recover the geometry of 3D parametric scenes using very few depth measurements from low-cost, commercially available time-of-flight sensors. These sensors offer very low spatial resolution (i.e., a single pixel), but image a wide field-of-view per pixel and capture detailed time-of-flight data in the form of time-resolved photon counts. This time-of-flight data encodes rich scene information and thus enables recovery of simple scenes from sparse measurements. We investigate the feasibility of using a distributed set of few measurements (e.g., as few as 15 pixels) to recover the geometry of simple parametric scenes with a strong prior, such as estimating the 6D pose of a known object. To achieve this, we design a method that utilizes both feed-forward prediction to infer scene parameters, and differentiable rendering within an analysis-by-synthesis framework to refine the scene parameter estimate. We develop hardware prototypes and demonstrate that our method effectively recovers object pose given an untextured 3D model in both simulations and controlled real-world captures, and show promising initial results for other parametric scenes. We additionally conduct experiments to explore the limits and capabilities of our imaging solution.
Carter Sifferman, Yiquan Li, Fangzhou Mu, Michael Gleicher, Mohit Gupta 0001, Yin Li 0003
ICCV7
2025 Learning to Inference Adaptively for Multimodal Large Language Models
abstract
Multimodal Large Language Models (MLLMs) have shown impressive capabilities in visual reasoning, yet come with substantial computational cost, limiting their deployment in resource-constrained settings. Despite recent effort on improving the efficiency of MLLMs, prior solutions fall short in responding to varying runtime conditions, in particular changing resource availability (e.g., contention due to the execution of other programs on the device). To bridge this gap, we introduce AdaLLaVA, an adaptive inference framework that learns to dynamically reconfigure operations in an MLLM during inference, accounting for the input data and a latency budget. We conduct extensive experiments across benchmarks involving question-answering, reasoning, and hallucination. Our results show that AdaLLaVA effectively adheres to input latency budget, achieving varying accuracy and latency tradeoffs at runtime. Further, we demonstrate that AdaLLaVA adapts to both input latency and content, can be integrated with token selection for enhanced efficiency, and generalizes across MLLMs. Our project webpage with code release is at https://zhuoyan-xu.github.io/ada-llava/.
Zhuoyan Xu, Khoi D. Nguyen 0001, Preeti Mukherjee, Saurabh Bagchi, Somali Chaterji, Yingyu Liang, Yin Li 0003
ICCV7
2025 AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning
abstract
Large language models (LLMs) have enabled the creation of multi-modal LLMs that exhibit strong comprehension of visual data such as images and videos. However, these models usually rely on extensive visual tokens from visual encoders, leading to high computational demands, which limits their applicability in resource-constrained environments and for long-context tasks. In this work, we propose a training-free adaptive inference method for multi-modal LLMs that can accommodate a broad range of efficiency requirements with a minimum performance drop. Our method consists of a) iterative token merging based on embedding similarity before LLMs, and b) progressive token pruning within LLM layers based on multi-modal importance. With a minimalist design, our method can be applied to both video and image LLMs. Extensive experiments on diverse video and image benchmarks demonstrate that our method substantially reduces computation load (e.g., a $\textbf{7-fold}$ reduction in FLOPs) while preserving the performance of video and image LLMs. Further, at a similar computational cost, our method outperforms the state-of-the-art methods in long video understanding (e.g., $\textbf{+4.6}$ on MLVU). Additionally, our in-depth analysis provides insights into token redundancy and LLM layer behaviors, offering guidance for future research in designing efficient multi-modal LLMs. Our code is available at https://github.com/LaVi-Lab/AIM.
Yiwu Zhong, Zhuoming Liu 0001, Yin Li 0003, Liwei Wang 0009
ICCV3
2025 LETS Forecast: Learning Embedology for Time Series Forecasting
abstract
Real-world time series are often governed by complex nonlinear dynamics. Understanding these underlying dynamics is crucial for precise future prediction. While deep learning has achieved major success in time series forecasting, many existing approaches do not explicitly model the dynamics. To bridge this gap, we introduce DeepEDM, a framework that integrates nonlinear dynamical systems modeling with deep neural networks. Inspired by empirical dynamic modeling (EDM) and rooted in Takens' theorem, DeepEDM presents a novel deep model that learns a latent space from time-delayed embeddings, and employs kernel regression to approximate the underlying dynamics, while leveraging efficient implementation of softmax attention and allowing for accurate prediction of future time steps. To evaluate our method, we conduct comprehensive experiments on synthetic data of nonlinear dynamical systems as well as real-world time series across domains. Our results show that DeepEDM is robust to input noise, and outperforms state-of-the-art methods in forecasting accuracy. Our code is available at: https://abrarmajeedi.github.io/deep_edm.
Abrar Majeedi, Gajjala Viswanatha Reddy, Satya Sai Srinath Namburi, Nada Magdi Elkordi, Yin Li 0003
ICML5
2025 Agile3D: Adaptive Contention- and Content-Aware 3D Object Detection for Embedded GPUs
abstract
Efficient 3D perception is critical for autonomous systems—self-driving vehicles, drones—to navigate safely in dynamic environments. Accurate 3D object detection from LiDAR data must handle irregular, high-volume point clouds, variable latency from contention and scene complexity, and tight embedded GPU constraints. Balancing accuracy and latency under dynamic conditions is crucial, yet existing frameworks like Chanakya [NeurIPS '23], LiteReconfig [EuroSys '22], and AdaScale [MLSys '19] struggle with the unique demands of 3D detection. We present Agile3D, the first adaptive 3D system integrating a cross-model Multi-branch Execution Framework (MEF) and a Contention- and Content-Aware Reinforcement Learning-based controller (CARL). CARL dynamically selects the optimal execution branch using five novel MEF control knobs: encoding format, spatial resolution, spatial encoding, 3D feature extractor, and detection head. CARL uses supervised training for stable initial policies, then Direct Preference Optimization (DPO) to finetune branch selection without hand-crafted rewards, presenting the first application of DPO to branch scheduling in 3D detection. Comprehensive evaluations show that Agile3D achieves state-of-the-art performance, maintaining high accuracy across varying hardware contention levels and 100-500 ms latency budgets. On NVIDIA Orin and Xavier GPUs, it consistently leads the Pareto frontier, outperforming existing methods for efficient 3D detection.
Pengcheng Wang 0001, Zhuoming Liu 0001, Shayok Bagchi, Ran Xu 0003, Saurabh Bagchi, Yin Li 0003, Somali Chaterji
MobiSys6
2025 Instant Video Models: Universal Adapters for Stabilizing Image-Based Networks
abstract
When applied sequentially to video, frame-based networks often exhibit temporal inconsistency—for example, outputs that flicker between frames. This problem is amplified when the network inputs contain time-varying corruptions. In this work, we introduce a general approach for adapting frame-based models for stable and robust inference on video. We describe a class of stability adapters that can be inserted into virtually any architecture and a resource-efficient training process that can be performed with a frozen base network. We introduce a unified conceptual framework for describing temporal stability and corruption robustness, centered on a proposed accuracy-stability-robustness loss. By analyzing the theoretical properties of this loss, we identify the conditions where it produces well-behaved stabilizer training. Our experiments validate our approach on several vision tasks including denoising (NAFNet), image enhancement (HDRNet), monocular depth (Depth Anything v2), and semantic segmentation (DeepLabv3+). Our method improves temporal stability and robustness against a range of image corruptions (including compression artifacts, noise, and adverse weather), while preserving or improving the quality of predictions.
Matthew Dutson, Nathan Labiosa, Yin Li 0003, Mohit Gupta 0001
NeurIPS3
2025 Towards video-based injury risk assessment: predicting lifting loads from body pose trajectories
abstract
Manual material handling tasks, such as lifting and lowering, are ubiquitous across industry sectors. Overexertion during these tasks is among the leading causes of workplace injuries. Previous studies have shown that lifting load is a key factor in determining the risk of injury. However, existing methods for measuring the lifting load often rely on manual measurements, sensor fusion, or other techniques that are difficult to scale in practice. In this study, we present a vision-based approach to automatically predict lifting load by analyzing human body pose trajectories extracted from video alone. Specifically, our method employs person detection, visual tracking, and human body pose estimation to extract pose trajectories and their kinematic features, which are then used to train a Transformer model for load prediction. To evaluate our method, we conducted a human subjects study of 19 participants performing various lifting and lowering tasks with varying postures. Our method achieved an average accuracy of 74.8% to distinguish between light vs. heavy objects, and an average accuracy of 50.8% to identify three levels of lifting loads (light, medium, heavy) across lifting and lowering tasks. These results demonstrate a first step towards computer vision based solutions for automatic, noninvasive, scalable injury risk assessment for manual material handling tasks.
Fangzhou Mu, Robert G. Radwin, Yin Li 0003
Mach. Vis. Appl.4
2025 Physics to the Rescue: Deep Non-Line-of-Sight Reconstruction for High-Speed Imaging
abstract
Computational approach to imaging around the corner, or non-line-of-sight (NLOS) imaging, is becoming a reality thanks to major advances in imaging hardware and reconstruction algorithms. A recent development towards practical NLOS imaging, (Nam et al. 2021) demonstrated a high-speed non-confocal imaging system that operates at 5Hz, 100x faster than the prior art. This enormous gain in acquisition rate, however, necessitates numerous approximations in light transport, breaking many existing NLOS reconstruction methods that assume an idealized image formation model. To bridge the gap, we present a novel deep model that incorporates the complementary physics priors of wave propagation and volume rendering into a neural network for high-quality and robust NLOS reconstruction. This orchestrated design regularizes the solution space by relaxing the image formation model, resulting in a deep model that generalizes well on real captures despite being exclusively trained on synthetic data. Further, we devise a unified learning framework that enables our model to be flexibly trained using diverse supervision signals, including target intensity images or even raw NLOS transient measurements. Once trained, our model renders both intensity and depth images at inference time in a single forward pass, capable of processing more than 5 captures per second on a high-end GPU. Through extensive qualitative and quantitative experiments, we show that our method outperforms prior physics and learning based approaches on both synthetic and real measurements. We anticipate that our method along with the fast capturing system will accelerate future development of NLOS imaging for real world applications that require high-speed imaging.
Fangzhou Mu, Sicheng Mo, Jiayong Peng, Xiaochun Liu, Ji Hyun Nam, Siddeshwar Raghavan, Andreas Velten, Yin Li 0003
IEEE Trans. Pattern Anal. Mach. Intell.8
2025 A Single-Camera Method for Estimating Lift Asymmetry Angles Using Deep Learning Computer Vision Algorithms
abstract
A computer vision (CV) method to automatically measure the revised NIOSH lifting equation asymmetry angle (A) from a single camera is described and tested. A laboratory study involving ten participants performing various lifts was used to estimateAin comparison to ground truth joint coordinates obtained using 3-D motion capture (MoCap). To address challenges, such as obstructed views and limitations in camera placement in real-world scenarios, the CV method utilized video-derived coordinates from a selected set of landmarks. A 2-D pose estimator (HR-Net) detected landmark coordinates in each video frame, and a 3-D algorithm (VideoPose3D) estimated the depth of each 2-D landmark by analyzing its trajectories. The mean absolute precision error for the CV method, compared to MoCap measurements using the same subset of landmarks for estimatingA, was 6.25° (SD = 10.19°, N = 360). The mean absolute accuracy error of the CV method, compared against conventional MoCap landmark markers was 9.45° (SD = 14.01°,N= 360).
Zhengyang Lou, Zitong Zhan, Yin Li 0003, Yu Hen Hu, Ming-Lun Lu, Dwight Werren, Robert G. Radwin
IEEE Trans. Hum. Mach. Syst.4
2024 FreeControl: Training-Free Spatial Control of Any Text-to-Image Diffusion Model with Any Condition
abstract
Recent approaches such as ControlNet [59] offer users fine-grained spatial control over text-to-image (T2I) diffusion models. However, auxiliary modules have to be trained for each spatial condition type, model architecture, and checkpoint, putting them at odds with the diverse intents and preferences a human designer would like to convey to the AI models during the content creation process. In this work, we present FreeControl, a training-free approach for controllable T2I generation that supports multiple conditions, architectures, and checkpoints simultaneously. Free Control enforces structure guidance to facilitate the global alignment with a guidance image, and appearance guidance to collect visual details from images generated without control. Extensive qualitative and quantitative experiments demonstrate the superior performance of Free Control across a variety of pre-trained T2I models. In particular, FreeControl enables convenient training-free control over many different architectures and checkpoints, allows the challenging input conditions on which most of the existing training-free methods fail, and achieves competitive synthesis quality compared to training-based approaches. Project page: https://genforce.github.io/freecontrol/.
Sicheng Mo, Fangzhou Mu, Kuan Heng Lin, Bochen Guan, Yin Li 0003, Bolei Zhou
CVPR6
2024 SnAG: Scalable and Accurate Video Grounding
abstract
Temporal grounding of text descriptions in videos is a central problem in vision-language learning and video understanding. Existing methods often prioritize accuracy over scalability - they have been optimized for grounding only a few text queries within short videos, and fail to scale up to long videos with hundreds of queries. In this paper, we study the effect of cross-modal fusion on the scalability of video grounding models. Our analysis establishes late fusion as a more cost-effective fusion scheme for long-form videos with many text queries. Moreover, it leads us to a novel, video-centric sampling scheme for efficient training. Based on these findings, we present SnAG, a simple baseline for scalable and accurate video grounding. Without bells and whistles, SnAG is 43% more accurate and$l.5\times$faster than CONE, a state of the art for long-form video grounding on the challenging MAD dataset, while achieving highly competitive results on short videos. Our code is available at https://github.com/fmu2/snag_release.
Fangzhou Mu, Sicheng Mo, Yin Li 0003
CVPR3
2024 Towards 3D Vision with Low-Cost Single-Photon Cameras
abstract
We present a method for reconstructing 3D shape of arbitrary Lambertian objects based on measurements by miniature, energy-efficient, low-cost single-photon cameras. These cameras, operating as time resolved image sensors, illuminate the scene with a very fast pulse of diffuse light and record the shape of that pulse as it returns back from the scene at a high temporal resolution. We propose to model this image formation process, account for its non-idealities, and adapt neural rendering to reconstruct 3D geometry from a set of spatially distributed sensors with known poses. We show that our approach can successfully recover complex 3D shapes from simulated data. We further demonstrate 3D object reconstruction from real-world captures, utilizing measurements from a commodity proximity sensor. Our work draws a connection between image-based modeling and active range scanning, and offers a step towards 3D vision with single-photon cameras. Our project webpage is at https://cpsiff.github.io/towards_3d_vision/.
Fangzhou Mu, Carter Sifferman, Sacha Jungerman, Yiquan Li, Mark Han, Michael Gleicher, Mohit Gupta 0001, Yin Li 0003
CVPR8
2024 RICA2: Rubric-Informed, Calibrated Assessment of Actions
Abrar Majeedi, Gajjala Viswanatha Reddy, Satya Sai Srinath Namburi, Yin Li 0003
ECCV (63)4
2024 Towards Few-Shot Adaptation of Foundation Models via Multitask Finetuning
abstract
Foundation models have emerged as a powerful tool for many AI problems. Despite the tremendous success of foundation models, effective adaptation to new tasks, particularly those with limited labels, remains an open question and lacks theoretical understanding. An emerging solution with recent success in vision and NLP involves finetuning a foundation model on a selection of relevant tasks, before its adaptation to a target task with limited labeled samples. In this paper, we study the theoretical justification of this multitask finetuning approach. Our theoretical analysis reveals that with a diverse set of related tasks, this multitask finetuning leads to reduced error in the target task, in comparison to directly adapting the same pretrained model. We quantify the relationship between finetuning tasks and target tasks by diversity and consistency metrics, and further propose a practical task selection algorithm. We substantiate our theoretical claims with extensive empirical evidence. Further, we present results affirming our task selection algorithm adeptly chooses related finetuning tasks, providing advantages to the model performance on target tasks. We believe our study shed new light on the effective adaptation of foundation models to new tasks that lack abundant labels. Our code is available at https://github.com/OliverXUZY/Foudation-Model_Multitask.
Zhuoyan Xu, Zhenmei Shi, Fangzhou Mu, Yin Li 0003, Yingyu Liang
ICLR5
2024 BioDrone: A Bionic Drone-Based Single Object Tracking Benchmark for Robust Vision
Xin Zhao 0012, Jing Zhang 0110, Yimin Hu, Rongshuai Liu, Haibin Ling, Yin Li 0003, Renshu Li, Jiadong Li
Int. J. Comput. Vis.8
2023 Learning Procedure-aware Video Representation from Instructional Videos and Their Narrations
abstract
The abundance of instructional videos and their narrations over the Internet offers an exciting avenue for understanding procedural activities. In this work, we propose to learn video representation that encodes both action steps and their temporal ordering, based on a large-scale dataset of web instructional videos and their narrations, without using human annotations. Our method jointly learns a video representation to encode individual step concepts, and a deep probabilistic model to capture both temporal dependencies and immense individual variations in the step ordering. We empirically demonstrate that learning temporal ordering not only enables new capabilities for procedure reasoning, but also reinforces the recognition of individual steps. Our model significantly advances the state-of-the-art results on step classification (+2.8%/+3.3% on COIN / EPIC-Kitchens) and step forecasting (+7.4% on COIN). Moreover, our model attains promising results in zero-shot inference for step classification and fore-casting, as well as in predicting diverse and plausible steps for incomplete procedures. Our code is available at https://github.com/facebookresearch/ProcedureVRL.
Yiwu Zhong, Licheng Yu, Shangwen Li, Xueting Yan, Yin Li 0003
CVPR6
2023 Eventful Transformers: Leveraging Temporal Redundancy in Vision Transformers
abstract
Vision Transformers achieve impressive accuracy across a range of visual recognition tasks. Unfortunately, their accuracy frequently comes with high computational costs. This is a particular issue in video recognition, where models are often applied repeatedly across frames or temporal chunks. In this work, we exploit temporal redundancy between subsequent inputs to reduce the cost of Transformers for video processing. We describe a method for identifying and re-processing only those tokens that have changed significantly over time. Our proposed family of models, Eventful Transformers, can be converted from existing Transformers (often without any re-training) and give adaptive control over the compute cost at runtime. We evaluate our method on large-scale datasets for video object detection (ImageNet VID) and action recognition (EPIC-Kitchens 100). Our approach leads to significant computational savings (on the order of 2-4x) with only minor reductions in accuracy.
Matthew Dutson, Yin Li 0003, Mohit Gupta 0001
ICCV2
2023 Learned Compressive Representations for Single-Photon 3D Imaging
abstract
Single-photon 3D cameras can record the time-of-arrival of billions of photons per second with picosecond accuracy. One common approach to summarize the photon data stream is to build a per-pixel timestamp histogram, resulting in a 3D histogram tensor that encodes distances along the time axis. As the spatio-temporal resolution of the histogram tensor increases, the in-pixel memory requirements and output data rates can quickly become impractical. To overcome this limitation, we propose a family of linear compressive representations of histogram tensors that can be computed efficiently, in an online fashion, as a matrix operation. We design practical lightweight compressive representations that are amenable to an in-pixel implementation and consider the spatio-temporal information of each timestamp. Furthermore, we implement our proposed framework as the first layer of a neural network, which enables the joint end-to-end optimization of the compressive representations and a downstream SPAD data processing model. We find that a well-designed compressive representation can reduce in-sensor memory and data rates up to 2 orders of magnitude without significantly reducing 3D imaging quality. Finally, we analyze the power consumption implications through an on-chip implementation.
Felipe Gutierrez-Barragan, Fangzhou Mu, Andrei Ardelean, Atul Ingle, Claudio Bruschini, Edoardo Charbon, Yin Li 0003, Mohit Gupta 0001, Andreas Velten
ICCV7
2023 Spike-Based Anytime Perception
abstract
In many emerging computer vision applications, it is critical to adhere to stringent latency and power constraints. The current neural network paradigm of frame-based, floating-point inference is often ill-suited to these resource-constrained applications. Spike-based perception – enabled by spiking neural networks (SNNs) – is one promising alternative. Unlike conventional neural networks (ANNs), spiking networks exhibit smooth tradeoffs between latency, power, and accuracy. SNNs are the archetype of an "anytime algorithm" whose accuracy improves smoothly over time. This property allows SNNs to adapt their computational investment in response to changing resource constraints. Unfortunately, mainstream algorithms for training SNNs (i.e., those based on ANN-to-SNN conversion) tend to produce models that are inefficient in practice. To mitigate this problem, we propose a set of principled optimizations that reduce latency and power consumption by 1–2 orders of magnitude in converted SNNs. These optimizations leverage a set of novel efficiency metrics designed for anytime algorithms. We also develop a state-of-the-art simulator, SaRNN, which can simulate SNNs using commodity GPU hardware and neuromorphic platforms. We hope that the proposed optimizations, metrics, and tools will facilitate the future development of spike-based vision systems.
Matthew Dutson, Yin Li 0003, Mohit Gupta 0001
WACV2
2023 In the Eye of the Beholder: Gaze and Actions in First Person Video
abstract
We address the task of jointly determining what a person is doing and where they are looking based on the analysis of video captured by a headworn camera. To facilitate our research, we first introduce the EGTEA Gaze+ dataset. Our dataset comes with videos, gaze tracking data, hand masks and action annotations, thereby providing the most comprehensive benchmark for First Person Vision (FPV). Moving beyond the dataset, we propose a novel deep model for joint gaze estimation and action recognition in FPV. Our method describes the participant's gaze as a probabilistic variable and models its distribution using stochastic units in a deep network. We further sample from these stochastic units, generating an attention map to guide the aggregation of visual features for action recognition. Our method is evaluated on our EGTEA Gaze+ dataset and achieves a performance level that exceeds the state-of-the-art by a significant margin. More importantly, we demonstrate that our model can be applied to larger scale FPV dataset-EPIC-Kitchens even without using gaze, offering new state-of-the-art results on FPV action recognition.
Yin Li 0003, Miao Liu 0007, James M. Rehg
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Weakly supervised foreground learning for weakly supervised localization and detection
Chen-Lin Zhang, Yin Li 0003, Jianxin Wu 0001
Pattern Recognit.2
2023 Virtuoso: Energy- and Latency-aware Streamlining of Streaming Videos on Systems-on-Chips
abstract
Efficient and adaptive computer vision systems have been proposed to make computer vision tasks, such as image classification and object detection, optimized for embedded or mobile devices. These solutions, quite recent in their origin, focus on optimizing the model (a deep neural network) or the system by designing an adaptive system with approximation knobs. Despite several recent efforts, we show that existing solutions suffer from two major drawbacks. First , while mobile devices or systems-on-chips usually come with limited resources including battery power, most systems do not consider the energy consumption of the models during inference. Second , they do not consider the interplay between the three metrics of interest in their configurations, namely, latency, accuracy, and energy. In this work, we propose an efficient and adaptive video object detection system— Virtuoso , which is jointly optimized for accuracy, energy efficiency, and latency. Underlying Virtuoso is a multi-branch execution kernel that is capable of running at different operating points in the accuracy-energy-latency axes, and a lightweight runtime scheduler to select the best fit execution branch to satisfy the user requirement. We position this work as a first step in understanding the suitability of various object detection kernels on embedded boards in the accuracy-latency-energy axes, opening the door for further development in solutions customized to embedded systems and for benchmarking such solutions. Virtuoso is able to achieve up to 286 FPS on the NVIDIA Jetson AGX Xavier board, which is up to 45× faster than the baseline EfficientDet D3 and 15× faster than the baseline EfficientDet D0. In addition, we also observe up to 97.2% energy reduction using Virtuoso compared to the baseline YOLO (v3)—a widely used object detector designed for mobiles. To fairly compare with Virtuoso , we benchmark 15 state-of-the-art or widely used protocols, including Faster R-CNN (FRCNN) [NeurIPS’15], YOLO v3 [CVPR’16], SSD [ECCV’16], EfficientDet [CVPR’20], SELSA [ICCV’19], MEGA [CVPR’20], REPP [IROS’20], FastAdapt [EMDL’21], and our in-house adaptive variants of FRCNN+, YOLO+, SSD+, and EfficientDet+ (our variants have enhanced efficiency for mobiles). With this comprehensive benchmark, Virtuoso has shown superiority to all the above protocols, leading the accuracy frontier at every efficiency level on NVIDIA Jetson mobile GPUs. Specifically, Virtuoso has achieved an accuracy of 63.9%, which is more than 10% higher than some of the popular object detection models, FRCNN at 51.1% and YOLO at 49.5%.
Jayoung Lee, Pengcheng Wang 0001, Ran Xu 0003, Venkat R. Dasari, Noah Weston, Yin Li 0003, Saurabh Bagchi, Somali Chaterji
ACM Trans. Design Autom. Electr. Syst.7
2022 3D Photo Stylization: Learning to Generate Stylized Novel Views from a Single Image
abstract
Visual content creation has spurred a soaring interest given its applications in mobile photography and AR / VR. Style transfer and single-image 3D photography as two representative tasks have so far evolved independently. In this paper, we make a connection between the two, and address the challenging task of 3D photo stylization - generating stylized novel views from a single image given an arbitrary style. Our key intuition is that style transfer and view synthesis have to be jointly modeled. To this end, we propose a deep model that learns geometry-aware content features for stylization from a point cloud representation of the scene, resulting in high-quality stylized images that are consistent across views. Further, we introduce a novel training protocol to enable the learning using only 2D images. We demonstrate the superiority of our method via extensive qualitative and quantitative studies, and showcase key applications of our method in light of the growing demand for 3D content creation from 2D image assets.11Project page: http://pages.es.wise.edu/-fmu/style3d
Fangzhou Mu, Jian Wang 0100, Yin Li 0003
CVPR4
2022 Smartadapt: Multi-branch Object Detection Framework for Videos on Mobiles
abstract
Several recent works seek to create lightweight deep net-works for video object detection on mobiles. We observe that many existing detectors, previously deemed computationally costly for mobiles, intrinsically support adaptive inference, and offer a multi-branch object detection frame-work (MBODF). Here, an MBODF is referred to as a so-lution that has many execution branches and one can dy-namically choose from among them at inference time to sat-isfy varying latency requirements (e.g. by varying resolution of an input frame). In this paper, we ask, and answer, the wide-ranging question across all MBODFs: How to expose the right set of execution branches and then how to sched-ule the optimal one at inference time? In addition, we un-cover the importance of making a content-aware decision on which branch to run, as the optimal one is conditioned on the video content. Finally, we explore a content-aware scheduler, an Oracle one, and then a practical one, leveraging various lightweight feature extractors. Our evaluation shows that layered on Faster R-CNN-based MBODF, compared to 7 baselines, our Smartadapt achieves a higher Pareto optimal curve in the accuracy-vs-latency space for the ILSVRC VID dataset.
Ran Xu 0003, Fangzhou Mu, Jayoung Lee, Preeti Mukherjee, Somali Chaterji, Saurabh Bagchi, Yin Li 0003
CVPR7
2022 RegionCLIP: Region-based Language-Image Pretraining
abstract
Contrastive language-image pretraining (CLIP) using image-text pairs has achieved impressive results on image classification in both zero-shot and transfer learning set-tings. However, we show that directly applying such mod-els to recognize image regions for object detection leads to unsatisfactory performance due to a major domain shift: CLIP was trained to match an image as a whole to a text de-scription, without capturing the fine-grained alignment be-tween image regions and text spans. To mitigate this issue, we propose a new method called RegionCLIP that signifi-cantly extends CLIP to learn region-level visual representations, thus enabling fine-grained alignment between image regions and textual concepts. Our method leverages a CLIP model to match image regions with template captions, and then pretrains our model to align these region-text pairs in the feature space. When transferring our pretrained model to the open-vocabulary object detection task, our method outperforms the state of the art by 3.8 AP50 and 2.2 AP for novel categories on COCO and LVIS datasets, respectively. Further, the learned region representations support zero-shot inference for object detection, showing promising results on both COCO and LVIS datasets. Our code is available at https://github.com/microsoft/RegionCLIP.
Yiwu Zhong, Pengchuan Zhang, Chunyuan Li, Noel Codella, Liunian Harold Li, Luowei Zhou, Xiyang Dai, Lu Yuan 0001, Yin Li 0003, Jianfeng Gao 0001
CVPR10
2022 Event Neural Networks
Matthew Dutson, Yin Li 0003, Mohit Gupta 0001
ECCV (11)2
2022 3D Scene Inference from Transient Histograms
Sacha Jungerman, Atul Ingle, Yin Li 0003, Mohit Gupta 0001
ECCV (7)3
2022 Egocentric Activity Recognition and Localization on a 3D Map
Miao Liu 0007, Lingni Ma, Kiran K. Somasundaram, Yin Li 0003, Kristen Grauman, James M. Rehg
ECCV (13)4
2022 ActionFormer: Localizing Moments of Actions with Transformers
Chen-Lin Zhang, Jianxin Wu 0001, Yin Li 0003
ECCV (4)3
2022 LiteReconfig: cost and content aware reconfiguration of video object detection systems for mobile GPUs
abstract
An adaptive video object detection system selects different execution paths at runtime, based on video content and available resources, so as to maximize accuracy under a target latency objective (e.g., 30 frames per second). Such a system is well suited to mobile devices with limited computing resources, and often running multiple contending applications. Existing solutions suffer from two major drawbacks. First, collecting feature values to decide on an execution branch is expensive. Second, there is a switching overhead for transitioning between branches and this overhead depends on the transition pair. LiteReconfig, an efficient and adaptive video object detection framework, addresses these challenges. LiteReconfig features a cost-benefit analyzer to decide which features to use, and which execution branch to run, at inference time. Furthermore, LiteReconfig has a content-aware accuracy prediction model, to select an execution branch tailored for frames in a video stream. We demonstrate that LiteReconfig achieves significantly improved accuracy under a set of varying latency objectives than existing systems, while maintaining up to 50 fps on an NVIDIA AGX Xavier board. Our code, with DOI, is available at https://doi.org/10.5281/zenodo.6345733.
Ran Xu 0003, Jayoung Lee, Pengcheng Wang 0001, Saurabh Bagchi, Yin Li 0003, Somali Chaterji
EuroSys5
2022 Robust Scene Inference under Noise-Blur Dual Corruptions
abstract
Scene inference under low-light is a challenging problem due to severe noise in the captured images. One way to reduce noise is to use longer exposure during the capture. However, in the presence of motion (scene or camera motion), longer exposures lead to motion blur, resulting in loss of image information. This creates a trade-off between these two kinds of image degradations: motion blur (due to long exposure) vs. noise (due to short exposure), also referred as a dual image corruption pair in this paper. With the rise of cameras capable of capturing multiple exposures of the same scene simultaneously, it is possible to overcome this trade-off. Our key observation is that although the amount and nature of degradation varies for these different image captures, the semantic content remains the same across all images. To this end, we propose a method to leverage these multi exposure captures for robust inference under low-light and motion. Our method builds on a feature consistency loss to encourage similar results from these individual captures, and uses the ensemble of their final predictions for robust visual recognition. We demonstrate the effectiveness of our approach on simulated images as well as real captures with multiple exposures, and across the tasks of object detection and image classification. Project: https://wisionlab.com/project/noiseblurdual
Bhavya Goyal, Jean-François Lalonde, Yin Li 0003, Mohit Gupta 0001
ICCP3
2021 Nyströmformer: A Nyström-based Algorithm for Approximating Self-Attention
abstract
Transformers have emerged as a powerful tool for a broad range of natural language processing tasks. A key component that drives the impressive performance of Transformers is the self-attention mechanism that encodes the influence or dependence of other tokens on each specific token. While beneficial, the quadratic complexity of self-attention on the input sequence length has limited its application to longer sequences - a topic being actively studied in the community. To address this limitation, we propose Nyströmformer - a model that exhibits favorable scalability as a function of sequence length. Our idea is based on adapting the Nyström method to approximate standard self-attention with O(n) complexity. The scalability of Nyströmformer enables application to longer sequences with thousands of tokens. We perform evaluations on multiple downstream tasks on the GLUE benchmark and IMDB reviews with standard sequence length, and find that our Nyströmformer performs comparably, or in a few cases, even slightly better, than standard self-attention. On longer sequence tasks in the Long Range Arena (LRA) benchmark, Nyströmformer performs favorably relative to other efficient self-attention methods. Our code is available at https://github.com/mlpen/Nystromformer.
Yunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan, Glenn Fung, Yin Li 0003
AAAI6
2021 Dual-Stream Multiple Instance Learning Network for Whole Slide Image Classification With Self-Supervised Contrastive Learning
abstract
We address the challenging problem of whole slide image (WSI) classification. WSIs have very high resolutions and usually lack localized annotations. WSI classification can be cast as a multiple instance learning (MIL) problem when only slide-level labels are available. We propose a MIL-based method for WSI classification and tumor detection that does not require localized annotations. Our method has three major components. First, we introduce a novel MIL aggregator that models the relations of the instances in a dual-stream architecture with trainable distance measurement. Second, since WSIs can produce large or unbalanced bags that hinder the training of MIL models, we propose to use self-supervised contrastive learning to extract good representations for MIL and alleviate the issue of prohibitive memory cost for large bags. Third, we adopt a pyramidal fusion mechanism for multiscale WSI features, and further improve the accuracy of classification and localization. Our model is evaluated on two representative WSI datasets. The classification accuracy of our model compares favorably to fully-supervised methods, with less than 2% accuracy gap across datasets. Our results also outperform all previous MIL-based methods. Additional benchmark results on standard MIL datasets further demonstrate the superior performance of our MIL aggregator on general MIL problems.
Bin Li 0064, Yin Li 0003, Kevin W. Eliceiri
CVPR2
2021 Improving Weakly Supervised Visual Grounding by Contrastive Knowledge Distillation
abstract
Weakly supervised phrase grounding aims at learning region-phrase correspondences using only image-sentence pairs. A major challenge thus lies in the missing links between image regions and sentence phrases during training. To address this challenge, we leverage a generic object detector at training time, and propose a contrastive learning framework that accounts for both region-phrase and image-sentence matching. Our core innovation is the learning of a region-phrase score function, based on which an image-sentence score function is further constructed. Importantly, our region-phrase score function is learned by distilling from soft matching scores between the detected object names and candidate phrases within an image-sentence pair, while the image-sentence score function is supervised by ground-truth image-sentence pairs. The design of such score functions removes the need of object detection at test time, thereby significantly reducing the inference cost. Without bells and whistles, our approach achieves state-of-the-art results on visual phrase grounding, surpassing previous methods that require expensive object detectors at test time.
Liwei Wang 0009, Jing Huang 0014, Yin Li 0003, Kun Xu 0005, Zhengyuan Yang, Dong Yu 0001
CVPR3
2021 A Simple Baseline for Weakly-Supervised Scene Graph Generation
abstract
We investigate the weakly-supervised scene graph generation, which is a challenging task since no correspondence of label and object is provided. The previous work regards such correspondence as a latent variable which is iteratively updated via nested optimization of the scene graph generation objective. However, we further reduce the complexity by decoupling it into an efficient first-order graph matching module optimized via contrastive learning to obtain such correspondence, which is used to train a standard scene graph generation model. The extensive experiments show that such a simple pipeline can significantly surpass the previous state-of-the-art by more than 30% on the Visual Genome dataset, both in terms of graph matching accuracy and scene graph quality. We believe this work serves as a strong baseline for future research. Code is available at https://github.com/jshi31/WS-SGG.
Jing Shi 0005, Yiwu Zhong, Ning Xu 0007, Yin Li 0003, Chenliang Xu
ICCV4
2021 Learning to Generate Scene Graph from Natural Language Supervision
abstract
Learning from image-text data has demonstrated recent success for many recognition tasks, yet is currently limited to visual features or individual visual concepts such as objects. In this paper, we propose one of the first methods that learn from image-sentence pairs to extract a graphical representation of localized objects and their relationships within an image, known as scene graph. To bridge the gap between images and texts, we leverage an off-the-shelf object detector to identify and localize object instances, match labels of detected regions to concepts parsed from captions, and thus create "pseudo" labels for learning scene graph. Further, we design a Transformer-based model to predict these "pseudo" labels via a masked token prediction task. Learning from only image-sentence pairs, our model achieves 30% relative gain over a latest method trained with human-annotated unlocalized scene graphs. Our model also shows strong results for weakly and fully supervised scene graph generation. In addition, we explore an open-vocabulary setting for detecting scene graphs, and present the first result for open-set scene graph generation.
Yiwu Zhong, Jing Shi 0005, Chenliang Xu, Yin Li 0003
ICCV5
2020 Attention Distillation for Learning Video Representations
Miao Liu 0007, Yun Zhang 0015, Yin Li 0003, James M. Rehg
BMVC4
2020 Interpretable and Accurate Fine-grained Recognition via Region Grouping
abstract
We present an interpretable deep model for fine-grained visual recognition. At the core of our method lies the integration of region-based part discovery and attribution within a deep neural network. Our model is trained using image-level object labels, and provides an interpretation of its results via the segmentation of object parts and the identification of their contributions towards classification. To facilitate the learning of object parts without direct supervision, we explore a simple prior of the occurrence of object parts. We demonstrate that this prior, when combined with our region-based part discovery and attribution, leads to an interpretable model that remains highly accurate. Our model is evaluated on major fine-grained recognition datasets, including CUB-200, CelebA and iNaturalist. Our results compares favourably to state-of-the-art methods on classification tasks, and outperforms previous approaches on the localization of object parts.
Zixuan Huang 0001, Yin Li 0003
CVPR2
2020 Forecasting Human-Object Interaction: Joint Prediction of Motor Attention and Actions in First Person Video
Miao Liu 0007, Siyu Tang 0001, Yin Li 0003, James M. Rehg
ECCV (1)3
2020 Comprehensive Image Captioning via Scene Graph Decomposition
Yiwu Zhong, Liwei Wang 0009, Jianshu Chen, Dong Yu 0001, Yin Li 0003
ECCV (14)5
2020 Gradients as Features for Deep Representation Learning
Fangzhou Mu, Yingyu Liang, Yin Li 0003
ICLR3
2020 ApproxDet: content and contention-aware approximate object detection for mobiles
abstract
Advanced video analytic systems, including scene classification and object detection, have seen widespread success in various domains such as smart cities and autonomous systems. With an evolution of heterogeneous client devices, there is incentive to move these heavy video analytics workloads from the cloud to mobile devices for low latency and real-time processing and to preserve user privacy. However, most video analytic systems are heavyweight and are trained offline with some pre-defined latency or accuracy requirements. This makes them unable to adapt at runtime in the face of three types of dynamism --- the input video characteristics change, the amount of compute resources available on the node changes due to co-located applications, and the user's latency-accuracy requirements change. In this paper we introduce ApproxDet, an adaptive video object detection framework for mobile devices to meet accuracy-latency requirements in the face of changing content and resource contention scenarios. To achieve this, we introduce a multi-branch object detection kernel, which incorporates a data-driven modeling approach on the performance metrics, and a latency SLA-driven scheduler to pick the best execution branch at runtime. We evaluate ApproxDet on a large benchmark video dataset and compare quantitatively to AdaScale and YOLOv3. We find that ApproxDet is able to adapt to a wide variety of contention and content characteristics and outshines all baselines, e.g., it achieves 52% lower latency and 11.1% higher accuracy over YOLOv3. Our software is open-sourced at https://github.com/purdue-dcsl/ApproxDet.
Ran Xu 0003, Chen-Lin Zhang, Pengcheng Wang 0001, Jayoung Lee, Subrata Mitra, Somali Chaterji, Yin Li 0003, Saurabh Bagchi
SenSys7
2019 Learning Two-Branch Neural Networks for Image-Text Matching Tasks
abstract
Image-language matching tasks have recently attracted a lot of attention in the computer vision field. These tasks include image-sentence matching, i.e., given an image query, retrieving relevant sentences and vice versa, and region-phrase matching or visual grounding, i.e., matching a phrase to relevant regions. This paper investigates two-branch neural networks for learning the similarity between these two data modalities. We propose two network structures that produce different output representations. The first one, referred to as an embedding network, learns an explicit shared latent embedding space with a maximum-margin ranking loss and novel neighborhood constraints. Compared to standard triplet sampling, we perform improved neighborhood sampling that takes neighborhood information into consideration while constructing mini-batches. The second network structure, referred to as a similarity network, fuses the two branches via element-wise product and is trained with regression loss to directly predict a similarity score. Extensive experiments show that our networks achieve high accuracies for phrase localization on the Flickr30K Entities dataset and for bi-directional image-sentence retrieval on Flickr30K and MSCOCO datasets.
Liwei Wang 0009, Yin Li 0003, Jing Huang 0014, Svetlana Lazebnik
IEEE Trans. Pattern Anal. Mach. Intell.2
2019 Focal Boundary Guided Salient Object Detection
abstract
The performance of salient object segmentation has been significantly advanced by using deep convolutional networks. However, these networks often produce blob-like saliency maps without accurate object boundaries. This is caused by the limited spatial resolution of their feature maps after multiple pooling operations, and might hinder downstream applications that require precise object shapes. To address this issue, we propose a novel deep model-Focal Boundary Guided (Focal- BG) network. Our model is designed to jointly learn to segment salient object masks and detect salient object boundaries. Our key idea is that additional knowledge about object boundaries can help to precisely identify the shape of the object. Moreover, our model incorporates a refinement pathway to refine the mask prediction, and makes use of the focal loss to facilitate the learning of the hard boundary pixels. To evaluate our model, we conduct extensive experiments. Our Focal-BG network consistently outperforms state-of-the-art methods on five major benchmarks. We provide a detailed analysis of these results and demonstrate that our joint modeling of salient object boundary and mask helps to better capture shape details, especially in the vicinity of object boundaries.
Yupei Wang, Xin Zhao 0012, Xuecai Hu, Yin Li 0003, Kaiqi Huang
IEEE Trans. Image Process.4
2019 Deep Crisp Boundaries: From Boundaries to Higher-Level Tasks
abstract
Edge detection has made significant progress with the help of deep convolutional networks (ConvNet). These ConvNet-based edge detectors have approached human level performance on standard benchmarks. We provide a systematical study of these detectors' outputs. We show that the detection results did not accurately localize edge pixels, which can be adversarial for tasks that require crisp edge inputs. As a remedy, we propose a novel refinement architecture to address the challenging problem of learning a crisp edge detector using ConvNet. Our method leverages a top-down backward refinement pathway, and progressively increases the resolution of feature maps to generate crisp edges. Our results achieve superior performance, surpassing human accuracy when using standard criteria on BSDS500, and largely outperforming the state-of-the-art methods when using more strict criteria. More importantly, we demonstrate the benefit of crisp edge maps for several important applications in computer vision, including optical flow estimation, object proposal generation, and semantic segmentation.
Yupei Wang, Xin Zhao 0012, Yin Li 0003, Kaiqi Huang
IEEE Trans. Image Process.3
2018 3D-RCNN: Instance-Level 3D Object Reconstruction via Render-and-Compare
abstract
We present a fast inverse-graphics framework for instance-level 3D scene understanding. We train a deep convolutional network that learns to map image regions to the full 3D shape and pose of all object instances in the image. Our method produces a compact 3D representation of the scene, which can be readily used for applications like autonomous driving. Many traditional 2D vision outputs, like instance segmentations and depth-maps, can be obtained by simply rendering our output 3D scene model. We exploit class-specific shape priors by learning a low dimensional shape-space from collections of CAD models. We present novel representations of shape and pose, that strive towards better 3D equivariance and generalization. In order to exploit rich supervisory signals in the form of 2D annotations like segmentation, we propose a differentiable Render-and-Compare loss that allows 3D shape and pose to be learned with 2D supervision. We evaluate our method on the challenging real-world datasets of Pascal3D+ and KITTI, where we achieve state-of-the-art results.
Abhijit Kundu, Yin Li 0003, James M. Rehg
CVPR2
2018 Compositional Learning for Human Object Interaction
Keizo Kato, Yin Li 0003, Abhinav Gupta 0001
ECCV (14)2
2018 In the Eye of Beholder: Joint Learning of Gaze and Actions in First Person Video
Yin Li 0003, Miao Liu 0007, James M. Rehg
ECCV (5)1
2018 Densely Cascaded Shadow Detection Network via Deeply Supervised Parallel Fusion
abstract
Shadow detection is an important and challenging problem in computer vision. Recently, single image shadow detection had achieved major progress with the development of deep convolutional networks. However, existing methods are still vulnerable to background clutters, and often fail to capture the global context of an input image. These global contextual and semantic cues are essential for accurately localizing the shadow regions. Moreover, rich spatial details are required to segment shadow regions with precise shape. To this end, this paper presents a novel model characterized by a deeply supervised parallel fusion (DSPF) network and a densely cascaded learning scheme. The DSPF network achieves a comprehensive fusion of global semantic cues and local spatial details by multiple stacked parallel fusion branches, which are learned in a deeply supervised manner. Moreover, the densely cascaded learning scheme is employed to refine the spatial details. Our method is evaluated on two widely used shadow detection benchmarks. Experimental results show that our method outperforms state-of-the-arts by a large margin.
Yupei Wang, Xin Zhao 0012, Yin Li 0003, Xuecai Hu, Kaiqi Huang
IJCAI3
2018 Beyond Grids: Learning Graph Representations for Visual Recognition
abstract
We propose learning graph representations from 2D feature maps for visual recognition. Our method draws inspiration from region based recognition, and learns to transform a 2D image into a graph structure. The vertices of the graph define clusters of pixels ("regions"), and the edges measure the similarity between these clusters in a feature space. Our method further learns to propagate information across all vertices on the graph, and is able to project the learned graph representation back into 2D grids. Our graph representation facilitates reasoning beyond regular grids and can capture long range dependencies among regions. We demonstrate that our model can be trained from end-to-end, and is easily integrated into existing networks. Finally, we evaluate our method on three challenging recognition tasks: semantic segmentation, object detection and object instance segmentation. For all tasks, our method outperforms state-of-the-art methods.
Yin Li 0003, Abhinav Gupta 0001
NeurIPS1
2018 Adaptive Discrete Hypergraph Matching
abstract
This paper addresses the problem of hypergraph matching using higher-order affinity information. We propose a solver that iteratively updates the solution in the discrete domain by linear assignment approximation. The proposed method is guaranteed to converge to a stationary discrete solution and avoids the annealing procedure and ad-hoc post binarization step that are required in several previous methods. Specifically, we start with a simple iterative discrete gradient assignment solver. This solver can be trapped in an -circle sequence under moderate conditions, where is the order of the graph matching problem. We then devise an adaptive relaxation mechanism to jump out this degenerating case and show that the resulting new path will converge to a fixed solution in the discrete domain. The proposed method is tested on both synthetic and real-world benchmarks. The experimental results corroborate the efficacy of our method.
Junchi Yan, Yin Li 0003, Guitao Cao
IEEE Trans. Cybern.3
2017 First-Person Action Decomposition and Zero-Shot Learning
abstract
In this work, we decompose a first-person action into verb and noun. We then study how the coupling of an action's constituent verb and noun affects the learners' ability to learn them separately and to combine them to perform recognition. We compare different information fusion methods on conventional action recognition and zero-shot learning, of which the latter is a strong indication of the feature's ability to capture one concept (verb/noun) and not be confounded by the other. To achieve the decoupling of verb/noun concepts, we extract features that are specialized for each of them. Specifically, we use improved dense trajectories and convolutional neural network activations. We show that by constructing specialized features for the decomposed concepts, our method succeeds in zero-shot learning. More surprisingly, it also outperforms previous results in conventional action recognition when the performance gaps of different features on verb/noun concepts are significant.
Yun C. Zhang, Yin Li 0003, James M. Rehg
WACV2
2016 Unsupervised Learning of Edges
abstract
Data-driven approaches for edge detection have proven effective and achieve top results on modern benchmarks. However, all current data-driven edge detectors require manual supervision for training in the form of hand-labeled region segments or object boundaries. Specifically, human annotators mark semantically meaningful edges which are subsequently used for training. Is this form of strong, highlevel supervision actually necessary to learn to accurately detect edges? In this work we present a simple yet effective approach for training edge detectors without human supervision. To this end we utilize motion, and more specifically, the only input to our method is noisy semi-dense matches between frames. We begin with only a rudimentary knowledge of edges (in the form of image gradients), and alternate between improving motion estimation and edge detection in turn. Using a large corpus of video data, we show that edge detectors trained using our unsupervised scheme approach the performance of the same methods trained with full supervision (within 3-5%). Finally, we show that when using a deep network for the edge detector, our approach provides a novel pre-training scheme for object detection.
Yin Li 0003, Manohar Paluri, James M. Rehg, Piotr Dollár
CVPR1
2016 Learning Deep Structure-Preserving Image-Text Embeddings
abstract
This paper proposes a method for learning joint embeddings of images and text using a two-branch neural network with multiple layers of linear projections followed by nonlinearities. The network is trained using a largemargin objective that combines cross-view ranking constraints with within-view neighborhood structure preservation constraints inspired by metric learning literature. Extensive experiments show that our approach gains significant improvements in accuracy for image-to-text and textto-image retrieval. Our method achieves new state-of-theart results on the Flickr30K and MSCOCO image-sentence datasets and shows promise on the new task of phrase localization on the Flickr30K Entities dataset.
Liwei Wang 0009, Yin Li 0003, Svetlana Lazebnik
CVPR2
2015 Delving into egocentric actions
abstract
We address the challenging problem of recognizing the camera wearer's actions from videos captured by an egocentric camera. Egocentric videos encode a rich set of signals regarding the camera wearer, including head movement, hand pose and gaze information. We propose to utilize these mid-level egocentric cues for egocentric action recognition. We present a novel set of egocentric features and show how they can be combined with motion and object features. The result is a compact representation with superior performance. In addition, we provide the first systematic evaluation of motion, object and egocentric cues in egocentric action recognition. Our benchmark leads to several surprising findings. These findings uncover the best practices for egocentric actions, with a significant performance boost over all previous state-of-the-art methods on three publicly available datasets.
Yin Li 0003, Zhefan Ye, James M. Rehg
CVPR1
2015 Gaze-enabled egocentric video summarization via constrained submodular maximization
abstract
With the proliferation of wearable cameras, the number of videos of users documenting their personal lives using such devices is rapidly increasing. Since such videos may span hours, there is an important need for mechanisms that represent the information content in a compact form (i.e., shorter videos which are more easily browsable/sharable). Motivated by these applications, this paper focuses on the problem of egocentric video summarization. Such videos are usually continuous with significant camera shake and other quality issues. Because of these reasons, there is growing consensus that direct application of standard video summarization tools to such data yields unsatisfactory performance. In this paper, we demonstrate that using gaze tracking information (such as fixation and saccade) significantly helps the summarization task. It allows meaningful comparison of different image frames and enables deriving personalized summaries (gaze provides a sense of the camera wearer's intent). We formulate a summarization model which captures common-sense properties of a good summary, and show that it can be solved as a submodular function maximization with partition matroid constraints, opening the door to a rich body of work from combinatorial optimization. We evaluate our approach on a new gaze-enabled egocentric video dataset (over 15 hours), which will be a valuable standalone resource.
Jia Xu 0011, Lopamudra Mukherjee, Yin Li 0003, Jamieson Warner, James M. Rehg
CVPR3
2014 The Secrets of Salient Object Segmentation
abstract
In this paper we provide an extensive evaluation of fixation prediction and salient object segmentation algorithms as well as statistics of major datasets. Our analysis identifies serious design flaws of existing salient object benchmarks, called the dataset design bias, by over emphasising the stereotypical concepts of saliency. The dataset design bias does not only create the discomforting disconnection between fixations and salient object segmentation, but also misleads the algorithm designing. Based on our analysis, we propose a new high quality dataset that offers both fixation and salient object segmentation ground-truth. With fixations and salient object being presented simultaneously, we are able to bridge the gap between fixations and salient objects, and propose a novel method for salient object segmentation. Finally, we report significant benchmark progress on 3 existing datasets of segmenting salient objects.
Yin Li 0003, Christof Koch, James M. Rehg, Alan L. Yuille
CVPR1
2014 Joint Semantic Segmentation and 3D Reconstruction from Monocular Video
Abhijit Kundu, Yin Li 0003, Frank Dellaert, Fuxin Li, James M. Rehg
ECCV (6)2
2014 Graduated Consistency-Regularized Optimization for Multi-graph Matching
Junchi Yan, Yin Li 0003, Wei Liu 0005, Hongyuan Zha, Xiaokang Yang 0001, Stephen M. Chu
ECCV (1)2
2013 Decoding Children's Social Behavior
abstract
We introduce a new problem domain for activity recognition: the analysis of children's social and communicative behaviors based on video and audio data. We specifically target interactions between children aged 1-2 years and an adult. Such interactions arise naturally in the diagnosis and treatment of developmental disorders such as autism. We introduce a new publicly-available dataset containing over 160 sessions of a 3-5 minute child-adult interaction. In each session, the adult examiner followed a semi-structured play interaction protocol which was designed to elicit a broad range of social behaviors. We identify the key technical challenges in analyzing these behaviors, and describe methods for decoding the interactions. We present experimental results that demonstrate the potential of the dataset to drive interesting research questions, and show preliminary results for multi-modal activity recognition.
James M. Rehg, Gregory D. Abowd, Agata Rozga, Mario Romero, Mark A. Clements, Stan Sclaroff, Irfan A. Essa, Opal Y. Ousley, Yin Li 0003, Chanho Kim, Hrishikesh Rao 0001, Jonathan C. Kim, Liliana Lo Presti, Jianming Zhang 0001, Denis Lantsman, Jonathan Bidwell, Zhefan Ye
CVPR9
2013 Learning to Predict Gaze in Egocentric Video
abstract
We present a model for gaze prediction in egocentric video by leveraging the implicit cues that exist in camera wearer's behaviors. Specifically, we compute the camera wearer's head motion and hand location from the video and combine them to estimate where the eyes look. We further model the dynamic behavior of the gaze, in particular fixations, as latent variables to improve the gaze prediction. Our gaze prediction results outperform the state-of-the-art algorithms by a large margin on publicly available egocentric vision datasets. In addition, we demonstrate that we get a significant performance boost in recognizing daily actions and segmenting foreground objects by plugging in our gaze predictions into state-of-the-art methods.
Yin Li 0003, Alireza Fathi, James M. Rehg
ICCV1
2012 Learning sparse covariance patterns for natural scenes
abstract
For scene classification, patch-level linear features do not always work as well as handcrafted features. In this paper, we present a new model to greatly improve the usefulness of linear features in classification by introducing co-variance patterns. We analyze their properties, discuss the fundamental importance, and present a generative model to properly utilize them. With this set of covariance information, in our framework, even the most naive linear features that originally lack the vital ability in classification become powerful. Experiments show that the performance of our new covariance model based on linear features is comparable with or even better than handcrafted features in scene classification.
Liwei Wang 0009, Yin Li 0003, Jiaya Jia, Jian Sun 0001, David P. Wipf, James M. Rehg
CVPR2
2012 Learning to Recognize Daily Actions Using Gaze
Alireza Fathi, Yin Li 0003, James M. Rehg
ECCV (1)2
2012 Detecting eye contact using wearable eye-tracking glasses
abstract
We describe a system for detecting moments of eye contact between an adult and a child, based on a single pair of gaze-tracking glasses which are worn by the adult. Our method utilizes commercial gaze tracking technology to determine the adult's point of gaze, and combines this with computer vision analysis of video of the child's face to determine their gaze direction. Eye contact is then detected as the event of simultaneous, mutual looking at faces by the dyad. We report encouraging findings from an initial implementation and evaluation of this approach.
Zhefan Ye, Yin Li 0003, Alireza Fathi, Yi Han 0005, Agata Rozga, Gregory D. Abowd, James M. Rehg
UbiComp2
2011 Robust facial feature points extraction in color images
Yue Zhou 0005, Yin Li 0003, Meilin Ge
Eng. Appl. Artif. Intell.2
2010 Optimum Subspace Learning and Error Correction for Tensors
Yin Li 0003, Junchi Yan, Yue Zhou 0005, Jie Yang 0002
ECCV (3)1
2010 Tensor error correction for corrupted values in visual data
abstract
The multi-channel image or the video clip has the natural form of tensor. The values of the tensor can be corrupted due to noise in the acquisition process. We consider the problem of recovering a tensor L of visual data from its corrupted observations X = L + S, where the corrupted entries S are unknown and unbounded, but are assumed to be sparse. Our work is built on the recent studies about the recovery of corrupted low-rank matrix via trace norm minimization. We extend the matrix case to the tensor case by the definition of tensor trace norm in [6]. Furthermore, the problem of tensor is formulated as a convex optimization, which is much harder than its matrix form. Thus, we develop a high quality algorithm to efficiently solve the problem. Our experiments show potential applications of our method and indicate a robust and reliable solution.
Yin Li 0003, Yue Zhou 0005, Junchi Yan, Jie Yang 0002, Xiangjian He
ICIP1
2010 Visual saliency detection via rank-sparsity decomposition
abstract
Saliency mechanism has been considered crucial in the human visual system and helpful to object detection and recognition. This paper addresses a novel feature-based model for visual saliency detection. It consists of two steps: first, using the learned overcomplete sparse bases to represent image patches; and then, estimating saliency information via direct low-rank and sparsity matrix decomposition. We compare our model with the previous methods on natural images. Experimental results show that our model performs competitively for visual saliency detection task, and suggest the potential application of matrix decomposition and convex optimization for image analysis.
Junchi Yan, Yin Li 0003, Zhibin Niu, Yuncai Liu
ICIP3
2010 An Optimization Based Framework for Human Pose Estimation
abstract
In computer vision community, human pose estimation and nonrigid shape recovery have evolved into different subfields. The state-of-the-art optimization techniques have been applied to the problem of deformable surface reconstruction successfully and recent methods in this area have focused on designing formulations that are easier to solve. In general, these techniques lay their success on the assumption that sufficient 2-D-3-D correspondences can be detected. By contrast, confronted with the similar ambiguity problem, many techniques for human pose estimation adopt stochastic searching or discriminative predictions, which allow for more generative image cues. However, the global optimization cannot be guaranteed via the stochastic methods; and discriminative techniques usually suffer from inaccuracy. In this letter, we absorb ideas from both domains and propose a unified approach for articulated human pose estimation. Specifically, we optimize the human pose to account for the discriminative pose prediction, bone length preservation in parallel with the point-topoint image observation. Moreover, the L2norm minimization is solved iteratively as a linear system with high computational efficiency.
Junchi Yan, Shuhan Shen, Yin Li 0003, Yuncai Liu
IEEE Signal Process. Lett.3
2009 Visual Saliency Based on Conditional Entropy
Yin Li 0003, Yue Zhou 0005, Junchi Yan, Zhibin Niu, Jie Yang 0002
ACCV (1)1
2009 An Accelerated Human Motion Tracking System Based on Voxel Reconstruction under Complex Environments
Junchi Yan, Yin Li 0003, EnLiang Zheng, Yuncai Liu
ACCV (2)2
2009 Incremental sparse saliency detection
abstract
By the guidance of attention, human visual system is able to locate objects of interest in complex scene. We propose a new visual saliency detection model for both image and video. Inspired by biological vision, saliency is defined locally. Lossy compression is adopted, where the saliency of a location is measured by the Incremental Coding Length(ICL). The ICL is computed by presenting the center patch as the sparsest linear representation of its surroundings. The final saliency map is generated by accumulating the coding length. The model is tested on both images and videos. The results indicate a reliable and robust saliency of our method.
Yin Li 0003, Yue Zhou 0005, Xiaochao Yang, Jie Yang 0002
ICIP1
2006 Flash matting
abstract
In this paper, we propose a novel approach to extract mattes using a pair of flash/no-flash images. Our approach, which we call flash matting , was inspired by the simple observation that the most noticeable difference between the flash and no-flash images is the foreground object if the background scene is sufficiently distant. We apply a new matting algorithm called joint Bayesian flash matting to robustly recover the matte from flash/no-flash images, even for scenes in which the foreground and the background are similar or the background is complex. Experimental results involving a variety of complex indoors and outdoors scenes show that it is easy to extract high-quality mattes using an off-the-shelf, flash-equipped camera. We also describe extensions to flash matting for handling more general scenes.
Jian Sun 0001, Yin Li 0003, Sing Bing Kang, Harry Shum
ACM Trans. Graph.2
2005 Symmetric Stereo Matching for Occlusion Handling
abstract
In this paper, we propose a symmetric stereo model to handle occlusion in dense two-frame stereo. Our occlusion reasoning is directly based on the visibility constraint that is more general than both ordering and uniqueness constraints used in previous work. The visibility constraint requires occlusion in one image and disparity in the other to be consistent. We embed the visibility constraint within an energy minimization framework, resulting in a symmetric stereo model that treats left and right images equally. An iterative optimization algorithm is used to approximate the minimum of the energy using belief propagation. Our stereo model can also incorporate segmentation as a soft constraint. Experimental results on the Middlebury stereo images show that our algorithm is state-of-the-art.
Jian Sun 0001, Yin Li 0003, Sing Bing Kang
CVPR (2)2
2005 Video object cut and paste
abstract
In this paper, we present a system for cutting a moving object out from a video clip. The cutout object sequence can be pasted onto another video or a background image. To achieve this, we first apply a new 3D graph cut based segmentation approach on the spatial-temporal video volume. Our algorithm partitions watershed presegmentation regions into foreground and background while preserving temporal coherence. Then, the initial segmentation result is refined locally. Given two frames in the video sequence, we specify two respective windows of interest which are then tracked using a bi-directional feature tracking algorithm. For each frame in between these two given frames, the segmentation in each tracked window is refined using a 2D graph cut that utilizes a local color model. Moreover, we provide brush tools for the user to control the object boundary precisely wherever needed. Based on the accurate binary segmentation result, we apply coherent matting to extract the alpha mattes and foreground colors of the object.
Yin Li 0003, Jian Sun 0001, Harry Shum
ACM Trans. Graph.1
2004 Stereo Reconstruction from Multiperspective Panoramas
Yin Li 0003, Harry Shum, Chi-Keung Tang, Richard Szeliski
IEEE Trans. Pattern Anal. Mach. Intell.1
2004 Lazy snapping
abstract
In this paper, we present Lazy Snapping , an interactive image cutout tool. Lazy Snapping separates coarse and fine scale processing, making object specification and detailed adjustment easy . Moreover, Lazy Snapping provides instant visual feedback, snapping the cutout contour to the true object boundary efficiently despite the presence of ambiguous or low contrast edges. Instant feedback is made possible by a novel image segmentation algorithm which combines graph cut with pre-computed over-segmentation. A set of intuitive user interface (UI) tools is designed and implemented to provide flexible control and editing for the users. Usability studies indicate that Lazy Snapping provides a better user experience and produces better segmentation results than the state-of-the-art interactive image cutout tool, Magnetic Lasso in Adobe Photoshop.
Yin Li 0003, Jian Sun 0001, Chi-Keung Tang, Harry Shum
ACM Trans. Graph.1
2004 Pop-up light field: An interactive image-based modeling and rendering system
abstract
In this article, we present an image-based modeling and rendering system, which we call pop-up light field , that models a sparse light field using a set of coherent layers . In our system, the user specifies how many coherent layers should be modeled or popped up according to the scene complexity. A coherent layer is defined as a collection of corresponding planar regions in the light field images. A coherent layer can be rendered free of aliasing all by itself, or against other background layers. To construct coherent layers, we introduce a Bayesian approach, coherence matting , to estimate alpha matting around segmented layer boundaries by incorporating a coherence prior in order to maintain coherence across images.We have developed an intuitive and easy-to-use user interface (UI) to facilitate pop-up light field construction. The key to our UI is the concept of human-in-the-loop where the user specifies where aliasing occurs in the rendered image. The user input is reflected in the input light field images where pop-up layers can be modified. The user feedback is instant through a hardware-accelerated real-time pop-up light field renderer. Experimental results demonstrate that our system is capable of rendering anti-aliased novel views from a sparse light field.
Harry Shum, Jian Sun 0001, Shuntaro Yamazaki, Yin Li 0003, Chi-Keung Tang
ACM Trans. Graph.4
2003 Rendering driven depth reconstruction
abstract
Previous work on image-based rendering suggests that there is a tradeoff between the number of images and the amount of geometry required for anti-aliased rendering. For instance, plenoptic sampling theory indicates that visually acceptable rendering can be achieved when the input images are undersampled, if sufficient depth information is available for all the pixels. In this paper, we propose a novel vision reconstruction approach, rendering-driven depth recovery, to recover the amount of geometry that is necessary for anti-aliased rendering. Our approach contrasts conventional stereo reconstruction in that we do not intend to accurately reconstruct the depth for each and every single pixel, leading to a very efficient reconstruction algorithm. Our algorithm uses a block-based multi-layer depth representation, and searches in the depth space based on the causality criterion, by detecting double images. Experiments show that rendering systems using our rendering driven depth recovery algorithm can synthesize satisfactory novel views efficiently by using 'just enough geometry' recovered from undersampled input images.
Yin Li 0003, Xin Tong 0001, Chi-Keung Tang, Harry Shum
ICASSP (4)1
2003 Large environment rendering using plenoptic primitives
abstract
One of the most difficult tasks in computer graphics is to enable virtual walkthroughs in very large and complicated environments that are photorealistic, seamless, and in real time. Current image-based rendering techniques, while capable of photorealism and interactive speeds, have failed in practice to extend to visualizations of such environments. We demonstrate an approach that defines a virtual walkthrough experience using plenoptic primitives (PPs). A PP can be any type of local visual experience: 360/spl deg/ static panorama, panoramic video (PV), lumigraph/light field representation, or concentric mosaics (CMs). By combining them judiciously, user experience can be authored with significantly reduced effort while maintaining high-quality user experience. We illustrate our technique on synthetic and real environments using PVs and CMs and show how the problem of achieving smooth transitions among PVs and CMs can be solved by using position-dependent local geometries.
Sing Bing Kang, Minsheng Wu, Yin Li 0003, Harry Shum
IEEE Trans. Circuits Syst. Video Technol.3
2001 Efficient Dense Depth Estimation from Dense Multiperspective Panoramas
Yin Li 0003, Chi-Keung Tang, Harry Shum
ICCV1