Jian Wang 0100

dblp:39/449-100 · DBLP profile ↗
← Back
56ranked-venue papers
8as first author
44since 2021 · last 2026
0000-0001-5266-3808ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 38 · 4 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 38 · 6 first-author · 30 since 2021Systems, architecture and hardware · 3 · 2 first-authorComputer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Cost-Effective Communication: An Auction-based Method for Language Agent Interaction
abstract
Multi-agent systems (MAS) built on large language models (LLMs) often suffer from inefficient ''free-for-all'' communication, leading to exponential token costs and low signal-to-noise ratios that hinder their practical deployment. We challenge the notion that more communication is always beneficial, hypothesizing instead that the core issue is the absence of resource rationality. We argue that "free'' communication, by ignoring the principle of scarcity, inherently breeds inefficiency and unnecessary expenses. To address this, we introduce the Dynamic Auction-based Language Agent (DALA), a novel framework that treats communication bandwidth as a scarce and tradable resource. Specifically, our DALA regards inter-agent communication as a centralized auction, where agents learn to bid for the opportunity to speak based on the predicted value density of their messages. Thus, our DALA intrinsically encourages agents to produce concise, informative messages while filtering out low-value communication. Extensive and comprehensive experiments demonstrate that our economically-driven DALA achieves new state-of-the-art performance across seven challenging reasoning benchmarks, including 84.32% on MMLU and a 91.21% pass@1 rate on HumanEval. Note that this is accomplished with remarkable efficiency, i.e., our DALA uses only 6.25 million tokens, a fraction of the resources consumed by current state-of-the-art methods on GSM8K. Further analysis reveals that our DALA cultivates the emergent skill of strategic silence, effectively adapting its communication strategies from verbosity to silence in a dynamic manner via resource constraints.
Yijia Fan, Jusheng Zhang, Kaitong Cai, Chengpei Tang, Jian Wang 0100, Keze Wang
AAAI6
2026 3DAlign-DAER: Dynamic Attention Policy and Efficient Retrieval Strategy for Fine-grained 3D-Text Alignment at Scale
abstract
Despite recent advancements in 3D-text cross-modal alignment, existing state-of-the-art methods still struggle to align fine-grained textual semantics with detailed geometric structures, and their alignment performance degrades significantly when scaling to large-scale 3D databases. To overcome this limitation, we introduce 3DAlign-DAER, a unified framework designed to align text and 3D geometry via the proposed dynamic attention policy and the efficient retrieval strategy, capturing subtle correspondences for diverse cross-modal retrieval and classification tasks. Specifically, during the training, our proposed dynamic attention policy (DAP) employs the Hierarchical Attention Fusion (HAF) module to represent the alignment as learnable fine-grained token-to-point attentions. To optimize these attentions across different tasks and geometric hierarchies, our DAP further exploits the Monte Carlo tree search to dynamically calibrate HAF attention weights via a hybrid reward signal and further enhances the alignment between textual descriptions and local 3D geometry. During the inference, our 3DAlign-DAER introduces an Efficient Retrieval Strategy (ERS) to leverage efficient hierarchical searching in the large-scale embedding spaces, outperforming traditional methods (eg, KNN) in accuracy and efficiency. Furthermore, to facilitate text-3D alignment research and train our 3DAlign-DAER, we construct Align3D-2M, a large-scale dataset featuring 2M text-3D pairs, to provide sufficient fine-grained cross-modal annotations. Extensive and comprehensive experiments demonstrate the superior performance of our 3DAlign-DAER on diverse benchmarks.
Yijia Fan, Jusheng Zhang, Kaitong Cai, Jian Wang 0100, Keze Wang
AAAI5
2026 Top-Down Semantic Refinement for Image Captioning
abstract
Large Vision-Language Models (VLMs) face an inherent contradiction in image captioning: their powerful single-step generation capabilities often lead to a myopic decision-making process. This makes it difficult to maintain global narrative coherence while capturing rich details, a limitation that is particularly pronounced in tasks that require multi-step and complex scene description. To overcome this fundamental challenge, we redefine image captioning as a goal-oriented hierarchical refinement planning problem, and further propose a novel framework, named Top-Down Semantic Refinement (TDSR), which models the generation process as a Markov Decision Process (MDP). However, planning within the vast state space of a VLM presents a significant computational hurdle. Our core contribution, therefore, is the design of a highly efficient Monte Carlo Tree Search (MCTS) algorithm tailored for VLMs. By incorporating a visual-guided parallel expansion and a lightweight value network, our TDSR reduces the call frequency to the expensive VLM by an order of magnitude without sacrificing planning quality. Furthermore, an adaptive early stopping mechanism dynamically matches computational overhead to the image's complexity. Extensive experiments on multiple benchmarks, including DetailCaps, COMPOSITIONCAP, and POPE, demonstrate that our TDSR, as a plug-and-play module, can significantly enhance the performance of existing VLMs (e.g., LLaVA-1.5, Qwen2.5-VL) by achieving state-of-the-art or highly competitive results in fine-grained description, compositional generalization, and hallucination suppression.
Jusheng Zhang, Kaitong Cai, Jian Wang 0100, Chengpei Tang, Keze Wang
AAAI4
2026 LLM-CAS: Dynamic Neuron Perturbation for Real-Time Hallucination Correction
abstract
Large language models (LLMs) often generate hallucinated content lacking factual or contextual grounding, hindering their reliability in critical applications. Traditional methods like supervised fine-tuning and reinforcement learning from human feedback are data-intensive and computationally expensive, while static parameter editing struggles with context-dependent errors and catastrophic forgetting. To overcome these limitations, we introduce LLM-CAS, a framework that formulates real-time hallucination correction as a hierarchical reinforcement learning (HRL) problem. LLM-CAS trains an agent to learn a sophisticated policy, dynamically selecting optimal, temporary neuron perturbations during inference based on the immediate context. This learned, policy-driven approach provides greater adaptability than prior dynamic methods that rely on heuristic or pre-defined adjustments. As a result, LLM-CAS achieves significant performance gains across various LLMs, improving accuracy by 10.98 percentage points on StoryCloze, 2.71 points on TriviaQA, and 2.06 points on TruthfulQA's MC1 score, thereby outperforming static methods like ITI and CAA, as well as the dynamic SADI framework. This context-aware, efficient approach promises enhanced reliability for LLMs in high-stakes domains, with future potential for multimodal extensions.
Jusheng Zhang, Ningyuan Liu, Yijia Fan, Qinglin Zeng, Kaitong Cai, Jian Wang 0100, Keze Wang
AAAI7
2026 Reinforcement Learning for Diffusion LLMs via Energy-Based Gibbs Alignment
abstract
Yijia Fan, Jing Yang, Mingyu Liu, Kaitong Cai, Jian Wang, Keze Wang, Jusheng Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yijia Fan, Kaitong Cai, Jian Wang 0100, Keze Wang, Jusheng Zhang
ACL (1)5
2026 Snapmoji: Instant Generation of Animatable Dual-Stylized Avatars
abstract
Despite the increasing popularity of avatar systems such as Snapchat Bitmojis, existing production avatar platforms face several limitations, such as a limited number of predefined assets, tedious customization processes, and inefficient rendering requirements. Addressing these shortcomings, we introduce Snapmoji, an avatar generation system that instantly creates 3D avatars, and enables customization in a process we call dual-stylization. Snapmoji first maps a selfie of a user to a primary avatar (e.g., Bitmoji style) using a new technique we name Gaussian Domain Adaptation (GDA), then applies a secondary style (e.g., skeleton, yarn, toy) to the primary avatar, all while preserving the user’s identity. The generated 3D avatars can then be rendered an animated on mobile devices at 30–40 FPS.
Eric Ming Chen, Di Liu 0003, Sizhuo Ma, Michael Vasilkovsky, Bing Zhou 0001, Wenzhou Wang, Jiahao Luo, Dimitris N. Metaxas, Vincent Sitzmann, Jian Wang 0100
WACV11
2026 Dance Like a Chicken: Low-Rank Stylization for Human Motion Diffusion
Haim Sawdayee, Chuan Guo 0002, Guy Tevet, Bing Zhou 0001, Jian Wang 0100, Amit Bermano
Comput. Graph. Forum5
2026 Dance Like a Chicken: Low-Rank Stylization for Human Motion Diffusion
abstract
Abstract Text‐to‐motion generative models span a wide range of 3D human actions but struggle with nuanced stylistic attributes such as a “Chicken” style. Due to the scarcity of style‐specific data, existing approaches pull the generative prior towards a reference style, which often results in out‐of‐distribution, low‐quality generations. In this work, we introduce LoRA‐MDM, a lightweight framework for motion stylization that generalizes to complex actions while maintaining editability. Our key insight is that adapting the generative prior to include the style, while preserving its overall distribution, is more effective than modifying each individual motion during generation. Building on this idea, LoRA‐MDM learns to adapt the prior to include the reference style using only a few samples. The style can then be used in the context of different textual prompts for generation. The low‐rank adaptation shifts the motion manifold in a semantically meaningful way, enabling realistic style infusion even for actions not present in the reference samples. Moreover, preserving the distribution structure enables advanced operations such as style blending and motion editing. We compare LoRA‐MDM to state‐of‐the‐art stylized motion generation methods and demonstrate a favorable balance between text fidelity and style consistency. Project page at https://haimsaw.github.io/LoRA‐MDM/
Haim Sawdayee, Chuan Guo 0002, Guy Tevet, Bing Zhou 0001, Jian Wang 0100, Amit Bermano
Comput. Graph. Forum5
2026 A Survey on Human Interaction Motion Generation
Kewei Sui, Anindita Ghosh, Inwoo Hwang, Bing Zhou 0001, Jian Wang 0100, Chuan Guo 0002
Int. J. Comput. Vis.5
2026 Velocity Disambiguation for Video Frame Interpolation
abstract
Existing video frame interpolation (VFI) methods blindly predict where each object is at a specific timestep $t$t ("time indexing"), which struggles to predict precise object movements. Given two images of a baseball, there are infinitely many possible trajectories: accelerating or decelerating, straight or curved. This often results in blurry frames as the method averages out these possibilities. Instead of forcing the network to learn this complicated time-to-location mapping implicitly together with predicting the frames, we provide the network with an explicit hint on how far the object has traveled between start and end frames, a novel approach termed "distance indexing". This method offers a clearer learning goal for models, reducing the uncertainty tied to object speeds. We further observed that, even with this extra guidance, objects can still be blurry especially when they are equally far from both input frames (i.e., halfway in-between), due to the directional ambiguity in long-range motion. To solve this, we propose an iterative reference-based estimation strategy that breaks down a long-range prediction into several short-range steps. When integrating our plug-and-play strategies into state-of-the-art learning-based models, they exhibit markedly sharper outputs and superior perceptual quality in arbitrary time interpolations, using a uniform distance indexing map in the same format as time indexing without requiring extra computation. Furthermore, we demonstrate that if additional latency is acceptable, a continuous map estimator can be employed to compute a pixel-wise dense distance indexing using multiple nearby frames. Combined with efficient multi-frame refinement, this extension can further disambiguate complex motion, thus enhancing performance both qualitatively and quantitatively. Additionally, the ability to manually specify distance indexing allows for independent temporal manipulation of each object, providing a novel tool for video editing tasks such as re-timing.
Zhihang Zhong, Wei Wang 0333, Xiao Sun 0001, Yu Qiao 0001, Gurunandan Krishnan, Sizhuo Ma, Jian Wang 0100
IEEE Trans. Pattern Anal. Mach. Intell.8
2025 Ego4o: Egocentric Human Motion Capture and Understanding from Multi-Modal Input
abstract
This work focuses on tracking and understanding human motion using consumer wearable devices, such as VR/AR headsets, smart glasses, cellphones, and smartwatches. These devices provide diverse, multi-modal sensor inputs, including egocentric images, and 1-3 sparse IMU sensors in varied combinations. Motion descriptions can also accompany these signals. The diverse input modalities and their intermittent availability pose challenges for consistent motion capture and understanding. In this work, we present Ego4o (o for omni), a new framework for simultaneous human motion capture and understanding from multi-modal egocentric inputs. This method maintains performance with partial inputs while achieving better results when multiple modalities are combined. First, the IMU sensor inputs, the optional egocentric image, and text description of human motion are encoded into the latent space of a motion VQ-VAE. Next, the latent vectors are sent to the VQ-VAE decoder and optimized to track human motion. When motion descriptions are unavailable, the latent vectors can be input into a multi-modal LLM to generate human motion descriptions, which can further enhance motion capture accuracy. Quantitative and qualitative evaluations demonstrate the effectiveness of our method in predicting accurate human motion and high-quality motion descriptions. Project page: https://jianwang-mpi.github.io/ego4o.
Jian Wang 0100, Rishabh Dabral, Diogo C. Luvizon, Zhe Cao 0003, Lingjie Liu, Thabo Beeler, Christian Theobalt
CVPR1
2025 FRAME: Floor-aligned Representation for Avatar Motion from Egocentric Video
abstract
Egocentric motion capture with a head-mounted body-facing stereo camera is crucial for VR and AR applications but presents significant challenges such as heavy occlusions and limited annotated real-world data. Existing methods rely on synthetic pretraining and struggle to generate smooth and accurate predictions in real-world settings, particularly for lower limbs. Our work addresses these limitations by introducing a lightweight VR-based data collection setup with on-board, real-time 6D pose tracking. Using this setup, we collected the most extensive real-world dataset for ego-facing ego-mounted cameras to date in size and motion variability. Effectively integrating this multimodal input –device pose and camera feeds –is challenging due to the differing characteristics of each data source. To address this, we propose FRAME, a simple yet effective architecture that combines device pose and camera feeds for state-of-the-art body pose prediction through geometrically sound multimodal integration and can run at 300 FPS on modern hardware. Lastly, we showcase a novel training strategy to enhance the model’s generalization capabilities. Our approach exploits the problem’s geometric properties, yielding high-quality motion capture free from common artifacts in prior work. Qualitative and quantitative evaluations, along with extensive comparisons, demonstrate the effectiveness of our method. Data, code, and CAD designs will be available at vcai.mpi-inf.mpg.de/projects/FRAME.
Andrea Boscolo Camiletto, Jian Wang 0100, Eduardo Alvarado, Rishabh Dabral, Thabo Beeler, Marc Habermann, Christian Theobalt
CVPR2
2025 DrDiff: Dynamic Routing Diffusion with Hierarchical Attention for Breaking the Efficiency-Quality Trade-off
abstract
This paper introduces DrDiff, a novel framework for long-text generation that overcomes the efficiency-quality trade-off through three core technologies.First, we design a dynamic expert scheduling mechanism that intelligently allocates computational resources during the diffusion process based on text complexity, enabling more efficient handling of text generation tasks of varying difficulty.Second, we introduce a Hierarchical Sparse Attention (HSA) mechanism that adaptively adjusts attention patterns according to a variety of input lengths, reducing computational complexity from O(n 2 ) to O(n) while maintaining model performance.Finally, we propose a Semantic Anchor States (SAS) module that combines with DPM-solver++ to reduce diffusion steps, significantly improving generation speed.Comprehensive experiments on various long-text generation benchmarks demonstrate the superiority of our DrDiff over the existing SOTA methods.
Jusheng Zhang, Yijia Fan, Kaitong Cai, Zimeng Huang, Jian Wang 0100, Chengpei Tang, Keze Wang
EMNLP6
2025 Exposure-Limited Image Enhancement with Generative Diffusion Prior
abstract
Many consumer cameras are equipped with 8-bit image sensors, which often struggle to capture scenes with a High Dynamic Range (HDR). This limitation can result in overexposed or underexposed regions, a loss of fine details due to low bit-depth compression, skewed color distributions, and noticeable noise in dark areas. Traditional Standard Dynamic Range (SDR) image enhancement methods typically focus on color mapping by expanding the color range and adjusting brightness. However, they often fail to restore details in dynamic range extremes, i.e. regions where pixel values approach the minimum or maximum limits. We define “exposure-limited image enhancement” as the process of enhancing images with large missing areas due to exposure issues within the SDR space, which differs from existing “mis-exposed image enhancement” methods primarily aimed at correcting color distributions. To enhance these exposure-limited images and overcome the limitations of current models, we propose a novel two-stage approach. In the first stage, we remap color and brightness to a suitable range while preserving existing details. In the second stage, we use a diffusion prior to generate content in severely overexposed or underexposed regions, which are otherwise lost during capture. Notably, this generative refinement module can also serve as a plug-and-play component alongside existing enhancement methods. Extensive experiments demonstrate that our method significantly improves image quality and detail, outperforming state-of-the-art techniques in dynamic range extremes. The project page is at https://Sagiri0208.github.io.
Baiang Li, Sizhuo Ma, Yanhong Zeng, Xiaogang Xu 0002, Youqing Fang, Zhao Zhang 0001, Jian Wang 0100, Kai Chen 0026
ICCP7
2025 DiffBody: Human Body Image Restoration with Generative Diffusion Prior
abstract
Human body image restoration is crucial for various applications but remains challenging due to the limitations of generative models: General image restoration methods built on generative models may generate unnatural textures, noticeable structural misalignments, and significant loss of fine details. To address these shortcomings, we present DiffBody, a novel human body-aware diffusion model that incorporates domain-specific knowledge to significantly enhance restoration quality. Our approach adopts a two-stage framework: (1) a multi-branch joint diffusion model generates preliminary priors, including normal and depth maps supported by a robust reconstruction pre-processing step; (2) a restoration stage refines the output using a body-prior ControlNet and a color adapter, ensuring structural accuracy and color consistency. Extensive quantitative evaluations, qualitative evaluations, and user studies validate the superior performance of DiffBody in producing perceptually high-quality human body restoration results. Code is available at https://github.com/yimingz1218/DiffBody.
Lionel Z. Wang, Sizhuo Ma, Xinjie Li 0002, Zhihang Zhong, Jian Wang 0100
ICCP7
2025 Scenemi: Motion In-Betweening for Modeling Human-Scene Interactions
abstract
Modeling human-scene interactions (HSI) is essential for understanding and simulating everyday human behaviors. Recent approaches utilizing generative modeling have made progress in this domain; however, they are limited in controllability and flexibility for real-world applications. To address these challenges, we propose reformulating the HSI modeling problem as Scene-aware Motion In-betweening - a more tractable and practical task. We introduce SceneMI, a framework that supports several practical applications, including keyframe-guided character animation in 3D scenes and enhancing the motion quality of imperfect HSI data. SceneMI employs dual scene descriptors to comprehensively encode global and local scene context. Furthermore, our framework leverages the inherent denoising nature of diffusion models to generalize on noisy keyframes. Experimental results demonstrate SceneMI's effectiveness in scene-aware keyframe in-betweening and generalization to the real-world GIMO dataset, where motions and scenes are acquired by noisy IMU sensors and smartphones. We further showcase SceneMI's applicability in HSI reconstruction from monocular videos.
Inwoo Hwang, Bing Zhou 0001, Young Min Kim 0001, Jian Wang 0100, Chuan Guo 0002
ICCV4
2025 Ponimator: Unfolding Interactive Pose for Versatile Human-Human Interaction Animation
Chuan Guo 0002, Bing Zhou 0001, Jian Wang 0100
ICCV4
2025 T2Bs: Text-to-Character Blendshapes via Video Generation
abstract
We present T2Bs, a framework for generating high-quality, animatable character head morphable models from text by combining static text-to-3D generation with video diffusion. Text-to-3D models produce detailed static geometry but lack motion synthesis, while video diffusion models generate motion with temporal and multi-view geometric inconsistencies. T2Bs bridges this gap by leveraging deformable 3D Gaussian splatting to align static 3D assets with video outputs. By constraining motion with static geometry and employing a view-dependent deformation MLP, T2Bs (i) outperforms existing 4D generation methods in accuracy and expressiveness while reducing video artifacts and view inconsistencies, and (ii) reconstructs smooth, coherent, fully registered 3D geometries designed to scale for building morphable models with diverse, realistic facial motions. This enables synthesizing expressive, animatable character heads that surpass current 4D generation techniques.
Jiahao Luo, Chaoyang Wang 0001, Michael Vasilkovsky, Vladislav Shakhrai, Di Liu 0003, Peiye Zhuang, Sergey Tulyakov, Peter Wonka, Hsin-Ying Lee 0001, Jian Wang 0100
ICCV11
2025 KABB: Knowledge-Aware Bayesian Bandits for Dynamic Expert Coordination in Multi-Agent Systems
abstract
As scaling large language models faces prohibitive costs, multi-agent systems emerge as a promising alternative, though challenged by static knowledge assumptions and coordination inefficiencies. We introduce Knowledge-Aware Bayesian Bandits (KABB), a novel framework that enhances multi-agent system coordination through semantic understanding and dynamic adaptation. The framework features three key innovations: a customized knowledge distance model for deep semantic understanding, a dual-adaptation mechanism for continuous expert optimization, and a knowledge-aware Thompson Sampling strategy for efficient expert selection. Extensive evaluation demonstrates KABB achieves an optimal cost-performance balance, maintaining high performance while keeping computational demands relatively low in multi-agent coordination.
Jusheng Zhang, Zimeng Huang, Yijia Fan, Ningyuan Liu, Zhuojie Yang, Jiawei Yao, Jian Wang 0100, Keze Wang
ICML8
2025 SnapMoGen: Human Motion Generation from Expressive Texts
abstract
Text-to-motion generation has experienced remarkable progress in recent years. However, current approaches remain limited to synthesizing motion from short or general text prompts, primarily due to dataset constraints. This limitation undermines fine-grained controllability and generalization to unseen prompts. In this paper, we introduce SnapMoGen, a new text-motion dataset featuring high-quality motion capture data paired with accurate, \textit{expressive} textual annotations. The dataset comprises 20K motion clips totaling 44 hours, accompanied by 122 detailed textual descriptions averaging 48 words per description (vs. 12 words of HumanML3D). Importantly, these motion clips preserve original temporal continuity as they were in long sequences, facilitating research in long-term motion generation and blending. We also improve upon previous generative masked modeling approaches. Our model, MoMask++, transforms motion into \textbf{multi-scale} token sequences that better exploit the token capacity, and learns to generate all tokens using a single generative masked transformer. MoMask++ achieves state-of-the-art performance on both HumanML3D and OmniMotion benchmarks. Additionally, we demonstrate the ability to process casual user prompts by employing an LLM to reformat inputs to align with the expressivity and narration style of SnapMoGen.
Chuan Guo 0002, Inwoo Hwang, Jian Wang 0100, Bing Zhou 0001
NeurIPS3
2025 CF-VLM: CounterFactual Vision-Language Fine-tuning
abstract
Recent advances in vision-language models (VLMs) have greatly improved cross-modal semantic understanding, yet significant limitations remain in fine-grained discrimination and deep causal reasoning tasks. Existing VLMs often rely on superficial statistical correlations, lacking the ability to capture the underlying causal logic between visual and textual content. To address this, we propose the **CounterFactual Vision-Language Fine-tuning Model (CF-VLM)**, a novel framework that enhances the causal reasoning capabilities of VLMs through the targeted use of counterfactual samples. CF-VLM introduces three complementary training objectives: maintaining foundational cross-modal alignment, reinforcing the uniqueness, and stability of factual scene representations against coherent counterfactuals, and sharpening the model’s sensitivity to minimal but critical causal edits. Extensive experiments demonstrate that CF-VLM consistently outperforms strong baselines and state-of-the-art methods on compositional reasoning and generalization benchmarks. Furthermore, it shows promise in mitigating visual hallucinations, indicating improved factual consistency. Our CF-VLM provides a robust foundation for deploying VLMs in high-stakes, real-world scenarios requiring reliable reasoning and interpretability.
Jusheng Zhang, Kaitong Cai, Yijia Fan, Jian Wang 0100, Keze Wang
NeurIPS4
2025 GAM-Agent: Game-Theoretic and Uncertainty-Aware Collaboration for Complex Visual Reasoning
abstract
We propose **GAM-Agent**, a game-theoretic multi-agent framework for enhancing vision-language reasoning. Unlike prior single-agent or monolithic models, GAM-Agent formulates the reasoning process as a non-zero-sum game between base agents—each specializing in visual perception subtasks—and a critical agent that verifies logic consistency and factual correctness. Agents communicate via structured claims, evidence, and uncertainty estimates. The framework introduces an uncertainty-aware controller to dynamically adjust agent collaboration, triggering multi-round debates when disagreement or ambiguity is detected. This process yields more robust and interpretable predictions. Experiments on four challenging benchmarks—MMMU, MMBench, MVBench, and V*Bench—demonstrate that GAM-Agent significantly improves performance across various VLM backbones. Notably, GAM-Agent boosts the accuracy of small-to-mid scale models (e.g., Qwen2.5-VL-7B, InternVL3-14B) by 5–6\%, and still enhances strong models like GPT-4o by up to 2–3\%. Our approach is modular, scalable, and generalizable, offering a path toward reliable and explainable multi-agent multimodal reasoning.
Jusheng Zhang, Yijia Fan, Haoyi Jiang, Wenhao Chai, Jian Wang 0100, Keze Wang
NeurIPS7
2025 3D-Agent: A Tri-Modal Multi-Agent Responsive Framework for Comprehensive 3D Object Annotation
abstract
Driven by the applications in autonomous driving, robotics, and augmented reality, 3D object annotation is a critical task compared to 2D annotation, such as spatial complexity, occlusion, and viewpoint inconsistency. The existing methods relying on single models often struggle with these issues. In this paper, we introduce Tri-MARF, a novel framework that integrates tri-modal inputs (i.e., 2D multi-view images, text descriptions, and 3D point clouds) with multi-agent collaboration to enhance the 3D annotation process. Our Tri-MARF consists of three specialized agents: a vision-language model agent that generates multi-view descriptions, an information aggregation agent that selects optimal descriptions, and a gating agent that aligns text descriptions with 3D geometries for more refined captioning. Extensive experiments on the Objaverse-LVIS, Objaverse-XL, and ABO datasets demonstrate the superiority of our Tri-MARF, which achieves a CLIPScore of 88.7 (compared to 78.6–82.4 for other SOTA methods), retrieval accuracy of 45.2/43.8 (ViLT R@5), and an impressive throughput of 12,000 objects per hour on a single NVIDIA A100 GPU.
Jusheng Zhang, Yijia Fan, Zimo Wen, Jian Wang 0100, Keze Wang
NeurIPS4
2025 4KAgent: Agentic Any Image to 4K Super-Resolution
abstract
We present 4KAgent, a unified agentic super-resolution generalist system designed to universally upscale any image to 4K resolution (and even higher, if applied iteratively). Our system can transform images from extremely low resolutions with severe degradations, for example, highly distorted inputs at $256\times 256$, into crystal-clear, photorealistic 4K outputs. 4KAgent comprises three core components: (1) Profiling, a module that customizes the 4KAgent pipeline based on bespoke use cases; (2) A Perception Agent, which leverages vision-language models alongside image quality assessment experts to analyze the input image and make a tailored restoration plan; and (3) A Restoration Agent, which executes the plan, following a recursive execution-reflection paradigm, guided by a quality-driven mixture-of-experts policy to select the optimal output for each step. Additionally, 4KAgent embeds a specialized face restoration pipeline, significantly enhancing facial details in portrait and selfie photos. We rigorously evaluate our 4KAgent across 11 distinct task categories encompassing a total of 26 diverse benchmarks, setting new state-of-the-art on a broad spectrum of imaging domains. Our evaluations cover natural images, portrait photos, AI-generated content, satellite imagery, fluorescence microscopy, and medical imaging like fundoscopy, ultrasound, and X-ray, demonstrating superior performance in terms of both perceptual (e.g., NIQE, MUSIQ) and fidelity (e.g., PSNR) metrics. By establishing a novel agentic paradigm for low-level vision tasks, we aim to catalyze broader interest and innovation within vision-centric autonomous agents across diverse research communities. We release all the code, models, and results at: https://4kagent.github.io.
Yushen Zuo, Qi Zheng 0004, Renjie Li 0003, Jian Wang 0100, Yide Zhang, Gengchen Mai, Lihong V. Wang, James Zou 0001, Ming-Hsuan Yang 0001, Zhengzhong Tu
NeurIPS6
2025 Towards 4D human video stylization
abstract
We present a first step towards 4D (3D space and time) human video stylization, which addresses style transfer, novel view synthesis, and human animation within a unified framework. While numerous video stylization methods have been developed, they are typically restricted to rendering images in specific viewpoints of the input video, lacking the capability to generalize to novel views and novel poses in dynamic scenes. To overcome these limitations, we leverage Neural Radiance Fields (NeRFs) to represent and stylize videos within a single framework. Our method involves simultaneously representing the human subject and the surrounding scene using two NeRFs. This dual representation facilitates the animation of human subjects across various poses and novel viewpoints. A key innovation is our introduction of a geometry-guided tri-plane representation, which significantly boosts the efficiency and robustness of the feature representation compared to direct tri-plane optimization. Stylization is performed within the NeRF rendered feature space, which can reduce the computational burden compared to applying style transformation to the feature vector of sampled points. Extensive experiments demonstrate that the proposed method strikes a superior balance between stylized textures and temporal coherence, surpassing existing approaches. Furthermore, our framework uniquely extends its capabilities to accommodate novel poses and viewpoints, making it a versatile tool for creative human video stylization. The source code and results will be available at this github site . The stylized videos are available in this Youtube video . • We present a 3D video stylization framework for novel views and human poses. • A geometric prior is introduced to enhance tri-plane feature learning efficiency. • Our method balances stylized textures and temporal coherence for dynamic scenes.
Xinxin Zuo, Fangzhou Mu, Jian Wang 0100, Ming-Hsuan Yang 0001
Comput. Vis. Image Underst.4
2025 Personalized Restoration via Dual-Pivot Tuning
abstract
Generative diffusion models can serve as priors, ensuring that image restoration solutions adhere to natural image manifolds. For facial images, however, personalized priors are essential to accurately reconstruct individual-specific facial features. We propose Dual-Pivot Tuning - a simple yet effective two-stage approach to personalize blind restoration systems while preserving general prior integrity. Our key observation is that for efficient personalization, the diffusion model should be tuned around a fixed textual pivot in the first step, while in the second step a guiding network should be tuned in a generic (non-personalized) manner, using the personalized diffusion model as a fixed "pivot". This approach ensures that personalization does not interfere with the restoration process, producing results with a natural appearance that show high fidelity to both identity and degraded image attributes. We conducted extensive experiments with images of widely recognized individuals, evaluating our approach both qualitatively and quantitatively against relevant baselines. Notably, our personalized prior not only achieves superior identity fidelity, but also outperforms state-of-the-art generic priors in terms of overall image quality. Project webpage is https://personalized-restoration.github.io/ and code is available at https://github.com/personalized-restoration/personalized-restoration.
Pradyumna Chari, Sizhuo Ma, Daniil Ostashev, Achuta Kadambi, Gurunandan Krishnan, Jian Wang 0100, Kfir Aberman
IEEE Trans. Image Process.6
2025 Privacy-Preserving Visual Localization With Event Cameras
abstract
We consider the problem of client-server localization, where edge device users communicate visual data with the service provider for locating oneself against a pre-built 3D map. This localization paradigm is a crucial component for location-based services in AR/VR or mobile applications, as it is not trivial to store large-scale 3D maps and process fast localization on resource-limited edge devices. Nevertheless, conventional client-server localization systems possess numerous challenges in computational efficiency, robustness, and privacy-preservation during data transmission. Our work aims to jointly solve these challenges with a localization pipeline based on event cameras. By using event cameras, our system consumes low energy and maintains small memory bandwidth. Then during localization, we propose applying event-to-image conversion and leverage mature image-based localization, which achieves robustness even in low-light or fast-moving scenes. To further enhance privacy protection, we introduce privacy protection techniques at two levels. Network level protection aims to hide the entire user's view in private scenes using a novel split inference approach, while sensor level protection aims to hide sensitive user details such as faces with light-weight filtering. Both methods involve small client-side computation and localization performance loss, while significantly mitigating the feeling of insecurity as revealed in our user study. We thus project our method to serve as a building block for practical location-based services using event cameras.
Young Min Kim 0001, Ramzi Zahreddine, Weston A. Welge, Gurunandan Krishnan, Sizhuo Ma, Jian Wang 0100
IEEE Trans. Image Process.7
2024 DSL-FIQA: Assessing Facial Image Quality via Dual-Set Degradation Learning and Landmark-Guided Transformer
abstract
Generic Face Image Quality Assessment (GFIQA) evalu-ates the perceptual quality of facial images, which is crucial in improving image restoration algorithms and selecting high-quality face images for downstream tasks. We present a novel transformer-based method for GFIQA, which is aided by two unique mechanisms. First, a “Dual-Set Degradation Representation Learning” (DSL) mechanism uses facial images with both synthetic and real degradations to decouple degradation from content, ensuring gen-eralizability to real-world scenarios. This self-supervised method learns degradation features on a global scale, pro-viding a robust alternative to conventional methods that use local patch information in degradation learning. Second, our transformer leverages facial landmarks to emphasize visually salient parts of a face image in evaluating its per-ceptual quality. We also introduce a balanced and diverse Comprehensive Generic Face IQA (CGFIQA-40k) dataset of 40K images carefully designed to overcome the biases, in particular the imbalances in skin tone and gender represen-tation, in existing datasets. Extensive analysis and evaluation demonstrate the robustness of our method, marking a significant improvement over prior methods.
Gurunandan Krishnan, Sy-Yen Kuo, Sizhuo Ma, Jian Wang 0100
CVPR6
2024 RobustSAM: Segment Anything Robustly on Degraded Images
abstract
Segment Anything Model (SAM) has emerged as a transformative approach in image segmentation, acclaimed for its robust zero-shot segmentation capabilities and flexible prompting system. Nonetheless, its performance is challenged by images with degraded quality. Addressing this limitation, we propose the Robust Segment Anything Model (RobustSAM), which enhances SAM's performance on low-quality images while preserving its promptability and zero-shot generalization. Our method leverages the pre-trained SAM model with only marginal parameter increments and computational requirements. The additional parameters of RobustSAM can be optimized within 30 hours on eight GPUs, demonstrating its feasibility and practicality for typical research laboratories. We also introduce the Robust-Seg dataset, a collection of 688K image-mask pairs with different degradations designed to train and evaluate our model optimally. Extensive experiments across various segmentation tasks and datasets confirm RobustSAM's superior performance, especially under zero-shot conditions, underscoring its potential for extensive real-world application. Additionally, our method has been shown to effectively improve the performance of SAM-based downstream tasks such as single image dehazing and deblurring.
Yu-Jiet Vong, Sy-Yen Kuo, Sizhuo Ma, Jian Wang 0100
CVPR5
2024 Holodepth: Programmable Depth-Varying Projection via Computer-Generated Holography
Dorian Chan, Matthew O'Toole, Sizhuo Ma, Jian Wang 0100
ECCV (61)4
2024 Delving Deep into Engagement Prediction of Short Videos
Dasong Li, Baili Lu, Hongsheng Li 0001, Sizhuo Ma, Gurunandan Krishnan, Jian Wang 0100
ECCV (54)7
2024 Clearer Frames, Anytime: Resolving Velocity Ambiguity in Video Frame Interpolation
Zhihang Zhong, Gurunandan Krishnan, Xiao Sun 0001, Yu Qiao 0001, Sizhuo Ma, Jian Wang 0100
ECCV (33)6
2024 DisCO: Portrait Distortion Correction with Perspective-Aware 3D GANs
Zhixiang Wang 0001, Yu-Lun Liu 0001, Jia-Bin Huang 0001, Shin'ichi Satoh 0001, Sizhuo Ma, Gurunandan Krishnan, Jian Wang 0100
Int. J. Comput. Vis.7
2024 Learning to Recover Spectral Reflectance From RGB Images
abstract
This paper tackles spectral reflectance recovery (SRR) from RGB images. Since capturing ground-truth spectral reflectance and camera spectral sensitivity are challenging and costly, most existing approaches are trained on synthetic images and utilize the same parameters for all unseen testing images, which are suboptimal especially when the trained models are tested on real images because they never exploit the internal information of the testing images. To address this issue, we adopt a self-supervised meta-auxiliary learning (MAXL) strategy that fine-tunes the well-trained network parameters with each testing image to combine external with internal information. To the best of our knowledge, this is the first work that successfully adapts the MAXL strategy to this problem. Instead of relying on naive end-to-end training, we also propose a novel architecture that integrates the physical relationship between the spectral reflectance and the corresponding RGB images into the network based on our mathematical analysis. Besides, since the spectral reflectance of a scene is independent to its illumination while the corresponding RGB images are not, we recover the spectral reflectance of a scene from its RGB images captured under multiple illuminations to further reduce the unknown. Qualitative and quantitative evaluations demonstrate the effectiveness of our proposed network and of the MAXL. Our code and data are available at https://github.com/Dong-Huo/SRR-MAXL.
Dong Huo, Jian Wang 0100, Yiming Qian, Yee-Hong Yang
IEEE Trans. Image Process.2
2024 Light Codes for Fast Two-Way Human-Centric Visual Communication
abstract
Visual codes, such as QR codes, are widely used in several applications for conveying information to users. However, user interactions based on spatial codes (e.g., displaying codes on phone screens for exchanging contact information) are often tedious, time consuming, and prone to errors due to image corruptions such as noise, blur, saturation, and perspective distortions. We propose Light Codes (LICO), a novel method for fast and fluid exchange of information among users. Light codes are based on transmitting and receiving temporal codes (instead of spatial) using compact and low-cost transceiver devices. The resulting approach enables seamless and near instantaneous exchange of short messages among users with minimal physical and cognitive effort. We design novel coding techniques, hardware prototypes, and applications that are optimized for human-centric communication, and facilitate fast and fluid user-to-user interactions in various challenging conditions, including a range of distances, motion, and ambient illumination. We evaluate the performance of the proposed methods both via quantitative analysis and user study based comparisons with several existing approaches including display-camera links, Bluetooth, and near-field communication, which show strong preference toward Light Codes in various real-world application scenarios.
Mohit Gupta 0001, Jian Wang 0100, Karl Bayer, Shree K. Nayar
ACM Trans. Graph.2
2024 Perspective-Aligned AR Mirror with Under-Display Camera
abstract
Augmented reality (AR) mirrors are novel displays that have great potential for commercial applications such as virtual apparel try-on. Typically the camera is placed beside the display, leading to distorted perspectives during user interaction. In this paper, we present a novel approach to address this problem by placing the camera behind a transparent display, thereby providing users with a perspective-aligned experience. Simply placing the camera behind the display can compromise image quality due to optical effects. We meticulously analyze the image formation process, and present an image restoration algorithm that benefits from physics-based data synthesis and network design. Our method significantly improves image quality and outperforms existing methods especially on the underexplored wire and backscatter artifacts. We then carefully design a full AR mirror system including display and camera selection, real-time processing pipeline, and mechanical design. Our user study demonstrates that the system is exceptionally well-received by users, highlighting its advantages over existing camera configurations not only as an AR mirror, but also for video conferencing. Our work represents a step forward in the development of AR mirrors, with potential applications in retail, cosmetics, fashion, etc. The image restoration dataset and code are available at https://perspective-armirror.github.io/.
Jian Wang 0100, Sizhuo Ma, Karl Bayer, Yi Zhang 0108, Peihao Wang, Bing Zhou 0001, Shree K. Nayar, Gurunandan Krishnan
ACM Trans. Graph.1
2023 Energy-Efficient Adaptive 3D Sensing
abstract
Active depth sensing achieves robust depth estimation but is usually limited by the sensing range. Naively increasing the optical power can improve sensing range but induces eye-safety concerns for many applications, including autonomous robots and augmented reality. In this paper, we propose an adaptive active depth sensor that jointly optimizes range, power consumption, and eye-safety. The main observation is that we need not project light patterns to the entire scene but only to small regions of interest where depth is necessary for the application and passive stereo depth estimation fails. We theoretically compare this adaptive sensing scheme with other sensing strategies, such as full-frame projection, line scanning, and point scanning. We show that, to achieve the same maximum sensing distance, the proposed method consumes the least power while having the shortest (best) eye-safety distance. We implement this adaptive sensing scheme with two hardware prototypes, one with a phase-only spatial light modulator (SLM) and the other with a micro-electro-mechanical (MEMS) mirror and diffractive optical elements (DOE). Experimental results validate the advantage of our method and demonstrate its capability of acquiring higher quality geometry adaptively. Please see our project website for video results and code: https://btilmon.github.io/e3d.html.
Brevin Tilmon, Zhanghao Sun, Sanjeev J. Koppal, Georgios Evangelidis 0002, Ramzi Zahreddine, Gurunandan Krishnan, Sizhuo Ma, Jian Wang 0100
CVPR9
2023 Be Real in Scale: Swing for True Scale in Dual Camera Mode
abstract
Many mobile AR apps that use the front-facing camera can benefit significantly from knowing the metric scale of the user’s face. However, the true scale of the face is hard to measure because monocular vision suffers from a fundamental ambiguity in scale. The methods based on prior knowledge about the scene either have a large error or are not easily accessible. In this paper, we propose a new method to measure the face scale by a simple user interaction: the user only needs to swing the phone to capture two selfies while using the recently popular Dual Camera mode. This mode allows simultaneous streaming of the front camera and the rear cameras and has become a key feature in many social apps. A computer vision method is applied to first estimate the absolute motion of the phone from the images captured by two rear cameras, and then calculate the point cloud of the face by triangulation. We develop a prototype mobile app to validate the proposed method. Our user study shows that the proposed method is favored compared to existing methods because of its high accuracy and ease of use. Our method can be built into Dual Camera mode and can enable a wide range of applications (e.g., virtual try-on for online shopping, true-scale 3D face modeling, gaze tracking, and face anti-spoofing) by introducing true scale to smartphone-based XR. The code is available at https://github.com/ruiyu0/Swing-for-True-Scale.
Rui Yu 0002, Jian Wang 0100, Sizhuo Ma, Sharon X. Huang, Gurunandan Krishnan
ISMAR2
2023 QfaR: Location-Guided Scanning of Visual Codes from Long Distances
abstract
Visual codes such as QR codes provide a low-cost and convenient communication channel between physical objects and mobile devices, but typically operate when the code and the device are in close physical proximity. We propose a system, called QfaR, which enables mobile devices to scan visual codes across long distances even where the image resolution of the visual codes is extremely low. QfaR is based on location-guided code scanning, where we utilize a crowd-sourced database of physical locations of codes. Our key observation is that if the approximate location of the codes and the user is known, the space of possible codes can be dramatically pruned down. Then, even if every "single bit" from the low-resolution code cannot be recovered, QfaR can still identify the visual code from the pruned list with high probability. By applying computer vision techniques, QfaR is also robust against challenging imaging conditions, such as tilt, motion blur, etc. Experimental results with common iOS and Android devices show that QfaR can significantly enhance distances at which codes can be scanned, e.g., 3.6cm-sized codes can be scanned at a distance of 7.5 meters, and 0.5m-sized codes at about 100 meters. QfaR has many potential applications, and beyond our diverse experiments, we also conduct a simple case study on its use for efficiently scanning QR code-based badges to estimate event attendance.
Sizhuo Ma, Jian Wang 0100, Wenzheng Chen, Suman Banerjee 0001, Mohit Gupta 0001, Shree K. Nayar
MobiCom2
2023 A Unified Conditional Framework for Diffusion-based Image Restoration
abstract
Diffusion Probabilistic Models (DPMs) have recently shown remarkable performance in image generation tasks, which are capable of generating highly realistic images. When adopting DPMs for image restoration tasks, the crucial aspect lies in how to integrate the conditional information to guide the DPMs to generate accurate and natural output, which has been largely overlooked in existing works. In this paper, we present a unified conditional framework based on diffusion models for image restoration. We leverage a lightweight UNet to predict initial guidance and the diffusion model to learn the residual of the guidance. By carefully designing the basic module and integration module for the diffusion model block, we integrate the guidance and other auxiliary conditional information into every block of the diffusion model to achieve spatially-adaptive generation conditioning. To handle high-resolution images, we propose a simple yet effective inter-step patch-splitting strategy to produce arbitrary-resolution images without grid artifacts. We evaluate our conditional framework on three challenging tasks: extreme low-light denoising, deblurring, and JPEG restoration, demonstrating its significant improvements in perceptual quality and the generalization to restoration tasks. The code will be released at https://zhangyi-3.github.io/project/UCDIR/.
Yi Zhang 0108, Xiaoyu Shi 0002, Dasong Li, Xiaogang Wang 0001, Jian Wang 0100, Hongsheng Li 0001
NeurIPS5
2023 Glass Segmentation With RGB-Thermal Image Pairs
abstract
This paper proposes a new glass segmentation method utilizing paired RGB and thermal images. Due to the large difference between the transmission property of visible light and that of the thermal energy through the glass where most glass is transparent to the visible light but opaque to thermal energy, glass regions of a scene are made more distinguishable with a pair of RGB and thermal images than solely with an RGB image. To exploit such a unique property, we propose a neural network architecture that effectively combines an RGB-thermal image pair with a new multi-modal fusion module based on attention, and integrate CNN and transformer to extract local features and non-local dependencies, respectively. As well, we have collected a new dataset containing 5551 RGB-thermal image pairs with ground-truth segmentation annotations. The qualitative and quantitative evaluations demonstrate the effectiveness of the proposed approach on fusing RGB and thermal data for glass segmentation. Our code and data are available at https://github.com/Dong-Huo/RGB-T-Glass-Segmentation.
Dong Huo, Jian Wang 0100, Yiming Qian, Yee-Hong Yang
IEEE Trans. Image Process.2
2022 3D Photo Stylization: Learning to Generate Stylized Novel Views from a Single Image
abstract
Visual content creation has spurred a soaring interest given its applications in mobile photography and AR / VR. Style transfer and single-image 3D photography as two representative tasks have so far evolved independently. In this paper, we make a connection between the two, and address the challenging task of 3D photo stylization - generating stylized novel views from a single image given an arbitrary style. Our key intuition is that style transfer and view synthesis have to be jointly modeled. To this end, we propose a deep model that learns geometry-aware content features for stylization from a point cloud representation of the scene, resulting in high-quality stylized images that are consistent across views. Further, we introduce a novel training protocol to enable the learning using only 2D images. We demonstrate the superiority of our method via extensive qualitative and quantitative studies, and showcase key applications of our method in light of the growing demand for 3D content creation from 2D image assets.11Project page: http://pages.es.wise.edu/-fmu/style3d
Fangzhou Mu, Jian Wang 0100, Yin Li 0003
CVPR2
2022 Seeing Far in the Dark with Patterned Flash
Zhanghao Sun, Jian Wang 0100, Shree K. Nayar
ECCV (6)2
2021 Seeing in Extra Darkness Using a Deep-Red Flash
abstract
We propose a new flash technique for low-light imaging, using deep-red light as an illuminating source. Our main observation is that in a dim environment, the human eye mainly uses rods for the perception of light, which are not sensitive to wavelengths longer than 620nm, yet the camera sensor still has a spectral response. We propose a novel modulation strategy when training a modern CNN model for guided image filtering, fusing a noisy RGB frame and a flash frame. This fusion network is further extended for video reconstruction. We have built a prototype with minor hardware adjustments and tested the new flash technique on a variety of static and dynamic scenes. The experimental results demonstrate that our method produces compelling reconstructions, even in extra dim conditions.
Jinhui Xiong, Jian Wang 0100, Wolfgang Heidrich, Shree K. Nayar
CVPR2
2020 Group Contextual Encoding for 3D Point Clouds
abstract
Global context is crucial for 3D point cloud scene understanding tasks. In this work, we extended the contextual encoding layer that was originally designed for 2D tasks to 3D Point Cloud scenarios. The encoding layer learns a set of code words in the feature space of the 3D point cloud to characterize the global semantic context, and then based on these code words, the method learns a global contextual descriptor to reweight the featuremaps accordingly. Moreover, compared to 2D scenarios, data sparsity becomes a major issue in 3D point cloud scenarios, and the performance of contextual encoding quickly saturates when the number of code words increases. To mitigate this problem, we further proposed a group contextual encoding method, which divides the channel into groups and then performs encoding on group-divided feature vectors. This method facilitates learning of global context in grouped subspace for 3D point clouds. We evaluate the effectiveness and generalizability of our method on three widely-studied 3D point cloud tasks. Experimental results have shown that the proposed method outperformed the VoteNet remarkably with 3 mAP on the benchmark of SUN-RGBD, with the metrics of mAP@ 0.25, and a much greater margin of 6.57 mAP on ScanNet with the metrics of mAP@ 0.5. Compared to the baseline of PointNet++, the proposed method leads to an accuracy of 86 %, outperforming the baseline by 1.5 %. Our proposed method have outperformed the non-grouping baseline methods across the board and establishes new state-of-the-art on these benchmarks.
Xu Liu 0017, Chengtao Li, Jian Wang 0100, Boxin Shi, Xiaodong He 0001
NeurIPS3
2019 Agile Depth Sensing Using Triangulation Light Curtains
abstract
Depth sensors like LIDARs and Kinect use a fixed depth acquisition strategy that is independent of the scene of interest. Due to the low spatial and temporal resolution of these sensors, this strategy can undersample parts of the scene that are important (small or fast moving objects), or oversample areas that are not informative for the task at hand (a fixed planar wall). In this paper, we present an approach and system to dynamically and adaptively sample the depths of a scene using the principle of triangulation light curtains. The approach directly detects the presence or absence of objects at specified 3D lines. These 3D lines can be sampled sparsely, non-uniformly, or densely only at specified regions. The depth sampling can be varied in real-time, enabling quick object discovery or detailed exploration of areas of interest. These results are achieved using a novel prototype light curtain system that is based on a 2D rolling shutter camera with higher light efficiency, working range, and faster adaptation than previous work, making it useful broadly for autonomous navigation and exploration.
Joseph R. Bartels, Jian Wang 0100, William Whittaker, Srinivasa G. Narasimhan
ICCV2
2019 Micro-Baseline Structured Light
abstract
We propose Micro-baseline Structured Light (MSL), a novel 3D imaging approach designed for small form-factor devices such as cell-phones and miniature robots. MSL operates with small projector-camera baseline and low-cost projection hardware, and can recover scene depths with computationally lightweight algorithms. The main observation is that a small baseline leads to small disparities, enabling a first-order approximation of the non-linear SL image formation model. This leads to the key theoretical result of the paper: the MSL equation, a linearized version of SL image formation. MSL equation is under-constrained due to two unknowns (depth and albedo) at each pixel, but can be efficiently solved using a local least squares approach. We analyze the performance of MSL in terms of various system parameters such as projected pattern and baseline, and provide guidelines for optimizing performance. Armed with these insights, we build a prototype to experimentally examine the theory and its practicality.
Vishwanath Saragadam, Raja Venkata, Jian Wang 0100, Shree K. Nayar, Mohit Gupta 0001
ICCV3
2018 Programmable Triangulation Light Curtains
Jian Wang 0100, Joseph R. Bartels, William Whittaker, Aswin C. Sankaranarayanan, Srinivasa G. Narasimhan
ECCV (3)1
2017 Compressive spectral anomaly detection
abstract
We propose a novel compressive imager for detecting anomalous spectral profiles in a scene. We model the background spectrum as a low-dimensional subspace while assuming the anomalies to form a spatially-sparse set of spectral profiles different from the background. Our core contributions are in the form of a two-stage sensing mechanism. In the first stage, we estimate the subspace for the background spectrum by acquiring spectral measurements at a few randomly-selected pixels. In the second stage, we acquire spatially-multiplexed spectral measurements of the scene. We remove the contributions of the background spectrum from the spatially-multiplexed measurements by projecting onto the complementary subspace of the background spectrum; the resulting measurements are of a sparse matrix that encodes the presence and spectra of anomalies, which can be recovered using a Multiple Measurement Vector formulation. Theoretical analysis and simulations show significant speed up in acquisition time over other anomaly detection techniques. A lab prototype based on a DMD and a visible spectrometer validates our proposed imager.
Vishwanath Saragadam, Jian Wang 0100, Xin Li 0001, Aswin C. Sankaranarayanan
ICCP2
2017 Reflectance Capture Using Univariate Sampling of BRDFs
abstract
We propose the use of a light-weight setup consisting of a collocated camera and light source – commonly found on mobile devices – to reconstruct surface normals and spatially-varying BRDFs of near-planar material samples. A collocated setup provides only a 1-D “univariate” sampling of a 3-D isotropic BRDF. We show that a univariate sampling is sufficient to estimate parameters of commonly used analytical BRDF models. Subsequently, we use a dictionary-based reflectance prior to derive a robust technique for per-pixel normal and BRDF estimation. We demonstrate real-world shape and capture, and its application to material editing and classification, using real data acquired using a mobile phone.
Zhuo Hui, Kalyan Sunkavalli, Joon-Young Lee, Sunil Hadap, Jian Wang 0100, Aswin C. Sankaranarayanan
ICCV5
2016 Dual Structured Light 3D Using a 1D Sensor
Jian Wang 0100, Aswin C. Sankaranarayanan, Mohit Gupta 0001, Srinivasa G. Narasimhan
ECCV (6)1
2015 LiSens- A Scalable Architecture for Video Compressive Sensing
abstract
The measurement rate of cameras that take spatially multiplexed measurements by using spatial light modulators (SLM) is often limited by the switching speed of the SLMs. This is especially true for single-pixel cameras where the photodetector operates at a rate that is many orders-of-magnitude greater than the SLM. We study the factors that determine the measurement rate for such spatial multiplexing cameras (SMC) and show that increasing the number of pixels in the device improves the measurement rate, but there is an optimum number of pixels (typically, few thousands) beyond which the measurement rate does not increase. This motivates the design of LiSens, a novel imaging architecture, that replaces the photodetector in the single-pixel camera with a 1D linear array or a line-sensor. We illustrate the optical architecture underlying LiSens, build a prototype, and demonstrate results of a range of indoor and outdoor scenes. LiSens delivers on the promise of SMCs: imaging at a megapixel resolution, at video rate, using an inexpensive low-resolution sensor.
Jian Wang 0100, Mohit Gupta 0001, Aswin C. Sankaranarayanan
ICCP1
2015 Photometric Stereo with Small Angular Variations
abstract
Most existing successful photometric stereo setups require large angular variations in illumination directions, which results in acquisition rigs that have large spatial extent. For many applications, especially involving mobile devices, it is important that the device be spatially compact. This naturally implies smaller angular variations in the illumination directions. This paper studies the effect of small angular variations in illumination directions to photometric stereo. We explore both theoretical justification and practical issues in the design of a compact and portable photometric stereo device on which a camera is surrounded by a ring of point light sources. We first derive the relationship between the estimation error of surface normal and the baseline of the point light sources. Armed with this theoretical insight, we develop a small baseline photometric stereo prototype to experimentally examine the theory and its practicality.
Jian Wang 0100, Yasuyuki Matsushita, Boxin Shi, Aswin C. Sankaranarayanan
ICCV1
2009 SRAM parametric failure analysis
abstract
With aggressive technology scaling, SRAM design has been seriously challenged by the difficulties in analyzing rare failure events. In this paper we propose to create statistical performance models with accuracy sufficient to facilitate probability extraction for SRAM parametric failures. A piecewise modeling technique is first proposed to capture the performance metrics over the large variation space. A controlled sampling scheme and a nested Monte Carlo analysis method are then applied for the failure probability extraction at cell-level and array-level respectively. Our 65nm SRAM example demonstrates that by combining the piecewise model and the fast probability extraction methods, we have significantly accelerated the SRAM failure analysis.
Jian Wang 0100, Soner Yaldiz, Xin Li 0001, Lawrence T. Pileggi
DAC1
2007 Parameterized Macromodeling for Analog System-Level Design Exploration
abstract
In this paper we propose a novel parameterized macromodeling technique for analog circuits. Unlike traditional macromodels that are only extracted for a small variation space, our proposed approach captures a significantly larger analog design space to facilitate system-level design exploration. Combining a novel piece-wise approximation algorithm and a new multi-point model-order-reduction approach, the proposed method generates compact macromodels covering the entire feasible design space. Our experiments demonstrate that using such models can achieve more than 60 x speed-up while incurring less than 4% overall error when varying design parameters by an order of magnitude.
Jian Wang 0100, Xin Li 0001, Lawrence T. Pileggi
DAC1
2005 Performance-centering optimization for system-level analog design exploration
abstract
In this paper we propose a novel analog design optimization methodology to address two key aspects of top-down system-level design: (1) how to optimally compare and select analog system architectures in the early phases of design; and (2) how to hierarchically propagate performance specifications from system level to circuit level to enable independent circuit block design. Importantly, due to the inaccuracy of early-stage system-level models, and the increasing magnitude of process and environmental variations, the system-level exploration must leave sufficient design margin to ensure a successful late-stage implementation. Therefore, instead of minimizing a design objective function, and thereby converging on a constraint boundary, we apply a novel performance centering optimization. Our proposed methodology centers the analog design in the performance space, and maximizes the distance to all constraint boundaries. We demonstrate that this early-stage design margin, which is measured by the volume of the inscribed ellipsoid lying inside the performance constraints, provides an excellent quality measure for comparing different system architectures. The efficacy of our performance centering approach is shown for analog design examples, including a complete clock data recovery system design and implementation.
Xin Li 0001, Jian Wang 0100, Lawrence T. Pileggi, Tun-Shih Chen, Wanju Chiang
ICCAD2