EDBT 2026 Demo / reviewers in the wild / expert
Fengyu Yang 0003
dblp:129/9492-3
· DBLP profile ↗
9ranked-venue papers
1as first author
9since 2021 · last 2025
0009-0001-1094-8204ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TextToucher: Fine-Grained Text-to-Touch GenerationabstractTactile sensation plays a crucial role in the development of multi-modal large models and embodied intelligence. To collect tactile data with minimal cost as possible, a series of studies have attempted to generate tactile images by vision-to-touch image translation. However, compared to text modality, visual modality-driven tactile generation cannot accurately depict human tactile sensation. In this work, we analyze the characteristics of tactile images in detail from two granularities: object-level (tactile texture, tactile shape), and sensor-level (gel status). We model these granularities of information through text descriptions and propose a fine-grained Text-to-Touch generation method (TextToucher) to generate high-quality tactile samples. Specifically, we introduce a multimodal large language model to build the text sentences about object-level tactile information and employ a set of learnable text prompts to represent the sensor-level tactile information. To better guide the tactile generation process with the built text information, we fuse the dual grains of text information and explore various dual-grain text conditioning methods within the diffusion transformer architecture. Furthermore, we propose a Contrastive Text-Touch Pre-training (CTTP) metric to precisely evaluate the quality of text-driven generated tactile data. Extensive experiments demonstrate the superiority of our TextToucher method. Jiahang Tu, Hao Fu 0023, Fengyu Yang 0003, Hanbin Zhao, Chao Zhang 0029, Hui Qian 0001 |
AAAI | 3 |
| 2025 | Visuo-Tactile Class-Incremental LearningabstractThe ability to associate sight with touch is essential for human and robot agents to understand material properties and to interact with the physical world. In the real-world scenarios, the robot agents often operate in a dynamically changing environment where new classes of objects are continually collected by visual and tactile sensors. In this article, we define this scenario as V isuo- T actile C lass- I ncremental L earning (VT-CIL). In practical VT-CIL, the robot needs to adapt to a new environment with constrained storage and computing resources, and suffers from the severe forgetting of vision and touch knowledge about old environments. To alleviate this problem, we consider visuo-tactile correlations in VT-CIL and propose a novel framework. It efficiently incorporates the Visuo-Tactile Cross-Modal Pseudo-Label-Consistent (VT-CMPLC) constraint, Dual-Visuo-Tactile Exemplars (DVT-E), and the Dual-Visuo-Tactile-Compatible (DVT-C) constraint. The old visual–tactile classes are preserved by the VT-CMPLC constraint and DVT-E, while the visuo-tactile correlations and the VT-CMPLC and DVT-E capabilities are enhanced by the DVT-C constraint. We built two benchmarks, the Touch-and-Go Class-Incremental (TaG-CI) benchmark and the ObjectFolder-Real Class-Incremental (OFR-CI) benchmark. Experimental results on TaG-CI and OFR-CI benchmarks demonstrate the effectiveness of our method against previous state-of-the-art class-incremental learning methods in VT-CIL. Hao Fu 0023, Fengyu Yang 0003, Boyang Wang 0008, Wei Ji 0008, Hanbin Zhao, Chao Zhang 0029, Roger Zimmermann, Hui Qian 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | APISR: Anime Production Inspired Real-World Anime Super-ResolutionabstractWhile real-world anime super-resolution (SR) has gained increasing attention in the SR community, existing methods still adopt techniques from the photorealistic domain. In this paper, we analyze the anime production workflow and rethink how to use characteristics of it for the sake of the real-world anime SR. First, we argue that video networks and datasets are not necessary for anime SR due to the repetition use of hand-drawing frames. Instead, we propose an anime image collection pipeline by choosing the least compressed and the most informative frames from the video sources. Based on this pipeline, we introduce the Anime Production-oriented Image (API) dataset. In addition, we identify two anime-specific challenges of distorted and faint hand-drawn lines and unwanted color artifacts. We address the first issue by introducing a prediction-oriented compression module in the image degradation model and a pseudo-ground truth preparation with enhanced hand-drawn lines. In addition, we introduce the balanced twin perceptual loss combining both anime and photorealistic high-level features to mitigate unwanted color artifacts and increase visual clarity. We evaluate our method through extensive experiments on the public benchmark, showing our method outperforms state-of-the-art anime dataset-trained approaches. The code is available at https://github.com/Kiteretsu77/APISR. Boyang Wang 0008, Fengyu Yang 0003, Xihang Yu, Chao Zhang 0029, Hanbin Zhao |
CVPR | 2 |
| 2024 | Binding Touch to Everything: Learning Unified Multimodal Tactile RepresentationsabstractTouch provides crucial information about the physical properties of the objects around us. Creating models that capture cross-modal associations between touch and other modalities, however, remains a challenging problem, due to wide variety of touch sensors and the intensive effort required to collect tactile data. We propose UniTouch, a unified model for vision-based touch sensors that connects their tactile signals to other modalities, including vision, language, and sound. We achieve this by aligning our tactile embeddings to pretrained image embeddings already associated with a variety of other modalities. We further propose learnable sensor-specific tokens, allowing the model to learn from a set of heterogeneous tactile sensors, all at the same time. UniTouch is capable of conducting various touch sensing tasks in a zero-shot setting, from robot grasping prediction to touch-based question answering. To the best of our knowledge, UniTouch is the first model to demonstrate these capabilities. Project Page: https://cfeng16.github.io/UniTouch/. Fengyu Yang 0003, Hyoungseob Park, Daniel Wang 0005, Yiming Dou, Ziyao Zeng, Xien Chen, Rit Gangopadhyay, Andrew Owens, Alex Wong 0001 |
CVPR | 1 |
| 2024 | WorDepth: Variational Language Prior for Monocular Depth EstimationabstractThree-dimensional (3D) reconstruction from a single image is an ill-posed problem with inherent ambiguities, i. e. scale. Predicting a 3D scene from text descriptionis) is similarly ill-posed, i. e. spatial arrangements of objects described. We investigate the question of whether two inher-ently ambiguous modalities can be used in conjunction to produce metric-scaled reconstructions. To test this, we fo-cus on monocular depth estimation, the problem of predicting a dense depth map from a single image, but with an additional text caption describing the scene. To this end, we begin by encoding the text caption as a mean and standard deviation; using a variational framework, we learn the distribution of the plausible metric reconstructions of 3D scenes corresponding to the text captions as a prior. To “select” a specific reconstruction or depth map, we encode the given image through a conditional sampler that samples from the latent space of the variational text encoder, which is then decoded to the output depth map. Our approach is trained alternatingly between the text and image branches: in one optimization step, we predict the mean and standard deviation from the text description and sample from a standard Gaussian, and in the other, we sample using a (image) conditional sampler. Once trained, we directly predict depth from the encoded text using the conditional sampler. We demonstrate our approach on indoor (NYUv2) and out-door (KITTI) scenarios, where we show that language can consistently improve performance in both. Code: https://github.com/Adonis-galaxy/WorDepth. Ziyao Zeng, Daniel Wang 0005, Fengyu Yang 0003, Hyoungseob Park, Stefano Soatto, Dong Lao, Alex Wong 0001 |
CVPR | 3 |
| 2024 | On the Viability of Monocular Depth Pre-training for Semantic Segmentation
Dong Lao, Fengyu Yang 0003, Daniel Wang 0005, Hyoungseob Park, Samuel Lu, Alex Wong 0001, Stefano Soatto |
ECCV (37) | 2 |
| 2024 | RSA: Resolving Scale Ambiguities in Monocular Depth Estimators through Language DescriptionsabstractWe propose a method for metric-scale monocular depth estimation. Inferring depth from a single image is an ill-posed problem due to the loss of scale from perspective projection during the image formation process. Any scale chosen is a bias, typically stemming from training on a dataset; hence, existing works have instead opted to use relative (normalized, inverse) depth. Our goal is to recover metric-scaled depth maps through a linear transformation. The crux of our method lies in the observation that certain objects (e.g., cars, trees, street signs) are typically found or associated with certain types of scenes (e.g., outdoor). We explore whether language descriptions can be used to transform relative depth predictions to those in metric scale. Our method, RSA , takes as input a text caption describing objects present in an image and outputs the parameters of a linear transformation which can be applied globally to a relative depth map to yield metric-scaled depth predictions. We demonstrate our method on recent general-purpose monocular depth models on indoors (NYUv2, VOID) and outdoors (KITTI). When trained on multiple datasets, RSA can serve as a general alignment module in zero-shot settings. Our method improves over common practices in aligning relative to metric depth and results in predictions that are comparable to an upper bound of fitting relative depth to ground truth via a linear transformation. Code is available at: https://github.com/Adonis-galaxy/RSA. Ziyao Zeng, Yangchao Wu, Hyoungseob Park, Daniel Wang 0005, Fengyu Yang 0003, Stefano Soatto, Dong Lao, Byung-Woo Hong, Alex Wong 0001 |
NeurIPS | 5 |
| 2024 | VCISR: Blind Single Image Super-Resolution with Video Compression Synthetic DataabstractIn the blind single image super-resolution (SISR) task, existing works have been successful in restoring image-level unknown degradations. However, when a single video frame becomes the input, these works usually fail to address degradations caused by video compression, such as mosquito noise, ringing, blockiness, and staircase noise. In this work, we for the first time, present a video compressionbased degradation model to synthesize low-resolution image data in the blind SISR task. Our proposed image synthesizing method is widely applicable to existing image datasets, so that a single degraded image can contain distortions caused by the lossy video compression algorithms. This overcomes the leak of feature diversity in video data and thus retains the training efficiency. By introducing video coding artifacts to SISR degradation models, neural networks can super-resolve images with the ability to restore video compression degradations, and achieve better results on restoring generic distortions caused by image compression as well. Our proposed approach achieves superior performance in SOTA no-reference Image Quality Assessment, and shows better visual quality on various datasets. In addition, we evaluate the SISR neural network trained with our degradation model on video super-resolution (VSR) datasets. Compared to architectures specifically designed for the VSR purpose, our method exhibits similar or better performance, evidencing that the presented strategy on infusing video-based degradation is generalizable to address more complicated compression artifacts even without temporal cues. The code is available at https://github.com/Kiteretsu77/VCISR-official. Boyang Wang 0008, Bowen Liu 0001, Fengyu Yang 0003 |
WACV | 4 |
| 2022 | RBC: Rectifying the Biased Context in Continual Semantic Segmentation
Hanbin Zhao, Fengyu Yang 0003, Xinghe Fu, Xi Li 0001 |
ECCV (34) | 2 |