VLDB 2026 Research / reviewers in the wild / expert
I-Sheng Fang
dblp:232/2119
· DBLP profile ↗
7ranked-venue papers
2as first author
6since 2021 · last 2026
0009-0001-8347-5938ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Text Slider: Efficient and Plug-and-Play Continuous Concept Control for Image/Video Synthesis via LoRA AdaptersabstractRecent advances in diffusion models have significantly improved image and video synthesis. In addition, several concept control methods have been proposed to enable fine-grained, continuous, and flexible control over free-form text prompts. However, these methods not only require intensive training time and GPU memory usage to learn the sliders or embeddings but also need to be retrained for different diffusion backbones, limiting their scalability and adaptability. To address these limitations, we introduce Text Slider, a lightweight, efficient and plug-and-play framework that identifies low-rank directions within a pre-trained text encoder, enabling continuous control of visual concepts while significantly reducing training time, GPU memory consumption, and the number of trainable parameters. Furthermore, Text Slider supports multi-concept composition and continuous control, enabling fine-grained and flexible manipulation in both image and video synthesis. We show that Text Slider enables smooth and continuous modulation of specific attributes while preserving the original spatial layout and structure of the input. Text Slider achieves significantly better efficiency: 5× faster training than Concept Slider and 47× faster than Attribute Control, while reducing GPU memory usage by nearly 2× and 4×, respectively. Project page: https://textslider.github.io Pin-Yen Chiu, I-Sheng Fang, Jun-Cheng Chen |
WACV | 2 |
| 2026 | KMOPS: Keypoint-Driven Method for Multi-Object Pose and Metric Size Estimation from Stereo ImagesabstractThe six-degree-of-freedom (6-DoF) pose and metric size estimation of multiple objects from RGB images alone remains a challenging task, particularly due to significant variations in object shape, appearance, and frequent occlusions in complex scenes. To address these challenges, we introduce KMOPS, a Keypoint-driven method tailored specifically for Multi-Object Pose and metric Size estimation from a single calibrated stereo image pair. Leveraging the stereo input, our approach first extracts the 2D keypoints of the enclosing bounding boxes of the objects across both views, and subsequently triangulates them to acquire metric 3D positions. Then, we obtain each object's rotation, translation, and dimensions by aligning the triangulated 3D keypoints to the canonical ones using a closed-form solution. Our formulation eliminates the need for predefined 3D search spaces or volumetric anchors, which are often required by other methods to constrain the vast 3D solution space. With extensive experiments on the challenging dataset Transparent Object Dataset (TOD) and StereOBJ-1M, we show that our method outperforms all competing methods with a simple and effective architecture. Ying-Kun Wu, Tzuhsuan Huang, I-Sheng Fang, Jun-Cheng Chen |
WACV | 4 |
| 2026 | US3Net: Ultralightweight Self-Supervised Stereo Matching Network using Depth-Aware Geometric Soft OcclusionabstractAbstract An ultralightweight self-supervised stereo matching network, called US $$^3$$ 3 Net, which requires only 12K parameters, is designed for efficient and accurate depth estimation using resource-constrained devices. US $$^3$$ 3 Net incorporates two key innovations: a low-complexity feature extraction module and a soft occlusion detection approach for performance improvement. First, we design a low-complexity feature extraction module to reduce the computational burden while preserving structural details necessary for stereo matching. By refining the encoder backbone and aggregation module, our design ensures a better balance between model complexity and accuracy. Second, to address occlusion-related errors in disparity estimation, we propose a novel occlusion detection method, called Depth-Aware Geometric Soft Occlusion (DAGSO), to adaptively define the occlusion confidence scores based on depth information. DAGSO can effectively mitigate false occlusions in distant regions and can improve the accuracy of disparity estimation. Experimental results using KITTI datasets demonstrate that US $$^3$$ 3 Net achieves state-of-the-art performance in terms of model complexity and depth estimation accuracy. It outperforms previous self-supervised stereo matching methods and monocular depth estimation methods in metrics such as AbsRel, SqRel, RMSE, and RMSElog at a reduction of parameter size by 47% compared with ES $$^3$$ 3 Net (23K). This makes US $$^3$$ 3 Net a practical solution for real-time depth estimation on edge devices such as drones and autonomous systems. Code is available at: https://github.com/g830319ag/US3Net. Po-Chung Jen, Tzu-Chi Liu, I-Sheng Fang, Hsiao-Chieh Wen, Chia-Lun Hsu, Ping-Yang Chen, Chang-Hsing Lee, Yong-Sheng Chen |
Int. J. Comput. Vis. | 3 |
| 2024 | Best of Both Sides: Integration of Absolute and Relative Depth Sensing Modalities Based on iToF and RGB Cameras
I-Sheng Fang, Walon Wei-Chen Chiu, Yong-Sheng Chen |
ICPR (16) | 1 |
| 2024 | Camera Settings as Tokens: Modeling Photography on Latent Diffusion Models
I-Sheng Fang, Yue-Hua Han, Jun-Cheng Chen |
SIGGRAPH Asia | 1 |
| 2022 | Single Image Reflection Removal Based on Knowledge-Distilling Content DisentanglementabstractWhen we shoot pictures through transparent media, such as glass, reflection can undesirably occur, obscuring the scene we intended to capture. Therefore, removing reflection is practical in image restoration. However, a reflective scene mixed with that behind the glass is challenging to be separated, considered significantly ill-posed. This letter addresses the single image reflection removal (SIRR) problem by proposing a knowledge-distilling-based content disentangling model that can effectively decompose the transmission and reflection layers. The experiments on benchmark SIRR datasets demonstrate that our method performs favorably against state-of-the-art SIRR methods. Yan-Tsung Peng, Kai-Han Cheng, I-Sheng Fang, Wen-Yi Peng, Jr-Shian Wu |
IEEE Signal Process. Lett. | 3 |
| 2020 | Self-Contained Stylization via Steganography for Reverse and Serial Style TransferabstractStyle transfer has been widely applied to give real-world images a new artistic look. However, given a stylized image, the attempts to use typical style transfer methods for de-stylization or transferring it again into another style usually lead to artifacts or undesired results. We realize that these issues are originated from the content inconsistency between the original image and its stylized output. Therefore, in this paper we advance to keep the content information of the input image during the process of style transfer by the power of steganography, with two approaches proposed: a two-stage model and an end-to-end model. We conduct extensive experiments to successfully verify the capacity of our models, in which both of them are able to not only generate stylized images of quality comparable with the ones produced by typical style transfer methods, but also effectively eliminate the artifacts introduced in reconstructing original input from a stylized image as well as performing multiple times of style transfer in series. Hung-Yu Chen, I-Sheng Fang, Chia-Ming Cheng, Walon Wei-Chen Chiu |
WACV | 2 |