EDBT 2026 Demo / reviewers in the wild / expert
Ta Ying Cheng
dblp:264/7281
· DBLP profile ↗
16ranked-venue papers
7as first author
14since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 7 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 6 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMsabstractMulti-view understanding, the ability to reconcile visual information across diverse viewpoints for effective navigation, manipulation, and 3D scene comprehension, is a fundamental challenge in Multi-Modal Large Language Models (MLLMs) to be used as embodied agents. While recent MLLMs have shown impressive advances in high-level reasoning and planning, they frequently fall short when confronted with multi-view geometric consistency and cross-view correspondence. To comprehensively evaluate the challenges of MLLMs in multi-view scene reasoning, we introduce All-Angles Bench, a human carefully benchmark with over 2,100 question-answer pairs from 90 diverse, real-world scenes. Our broad evaluation across 38 general-purpose and 3D spatial reasoning MLLMs reveals a substantial performance gap compared to humans. More critically, our analysis identifies two root failure modes: (1) cross-view object mismatch—the inability to establish consistent object correspondence across views; and (2) cross-view spatial misalignment—the failure to infer accurate camera poses and spatial layouts. These findings underscore a lack of multi-view awareness in current MLLMs, calling for architectural innovations beyond prompt tuning alone. We believe that our benchmark offers valuable insights toward building spatially-intelligent MLLMs. Chun-Hsiao Yeh, Shengbang Tong, Ta Ying Cheng, Ruoyu Wang 0014, Tianzhe Chu, Yuexiang Zhai, Yubei Chen, Shenghua Gao, Yi Ma 0001 |
AAAI | 4 |
| 2025 | MARBLE: Material Recomposition and Blending in CLIP-SpaceabstractEditing materials of objects in images based on exemplar images is an active area of research in computer vision and graphics. We propose MARBLE, a method for performing material blending and recomposing fine-grained material properties by finding material embeddings in CLIP-space and using that to control pre-trained text-to-image models. We improve exemplar-based material editing by finding a block in the denoising UNet responsible for material attribution. Given two material exemplar-images, we find directions in the CLIP-space for blending the materials. Further, we can achieve parametric control over fine-grained material attributes such as roughness, metallic, transparency, and glow using a shallow network to predict the direction for the desired material attribute change. We perform qualitative and quantitative analysis to demonstrate the efficacy of our proposed method. We also present the ability of our method to perform multiple edits in a single forward pass and applicability to painting. Ta Ying Cheng, Prafull Sharma, Mark Boss, Varun Jampani |
CVPR | 1 |
| 2024 | Learning Continuous 3D Words for Text-to-Image GenerationabstractCurrent controls over diffusion models (e.g., through text or ControlNet) for image generation fall short in recognizing abstract, continuous attributes like illumination direction or non-rigid shape change. In this paper, we present an approach for allowing users of text-to-image models to have fine-grained control of several attributes in an image. We do this by engineering special sets of input tokens that can be transformed in a continuous manner – we call them Continuous 3D Words. These attributes can, for example, be represented as sliders and applied jointly with text prompts for fine-grained control over image generation. Given only a single mesh and a rendering engine, we show that our approach can be adopted to provide continuous user control over several 3D-aware attributes, including time-of-day illumination, bird wing orientation, dollyzoom effect, and object poses. Our method is capable of conditioning image creation with multiple Continuous 3D Words and text descriptions simultaneously while adding no overhead to the generative process. Project Page: https://ttchengab.github.io/continuous_3d_words Ta Ying Cheng, Matheus Gadelha, Thibault Groueix, Matthew Fisher, Radomír Mech, Andrew Markham, Agathoniki Trigoni |
CVPR | 1 |
| 2024 | ZeST: Zero-Shot Material Transfer from a Single Image
Ta Ying Cheng, Prafull Sharma, Andrew Markham, Agathoniki Trigoni, Varun Jampani |
ECCV (1) | 1 |
| 2024 | SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D PriorsabstractCurrent state-of-the-art spatial reasoning-enhanced VLMs are trained to excel at spatial visual question answering (VQA). However, we believe that higher-level 3D-aware tasks, such as articulating dynamic scene changes and motion planning, require a fundamental and explicit 3D understanding beyond current spatial VQA datasets. In this work, we present SpatialPIN, a framework designed to enhance the spatial reasoning capabilities of VLMs through prompting and interacting with priors from multiple 3D foundation models in a zero-shot, training-free manner. Extensive experiments demonstrate that our spatial reasoning-imbued VLM performs well on various forms of spatial VQA and can extend to help in various downstream robotics tasks such as pick and stack and trajectory planning. Kai Lu 0003, Ta Ying Cheng, Agathoniki Trigoni, Andrew Markham |
NeurIPS | 3 |
| 2024 | Towards Learning Group-Equivariant Features for Domain Adaptive 3D DetectionabstractThe performance of 3D object detection in large outdoor point clouds deteriorates significantly in an unseen environment due to the inter-domain gap. To address these challenges, most existing methods for domain adaptation harness self-training schemes and attempt to bridge the gap by focusing on a single factor that causes the inter-domain gap, such as objects' sizes, shapes, and foreground density variation. However, the resulting adaptations suggest that there is still a substantial inter-domain gap left to be minimized. We argue that this is due to two limitations: 1) Biased pseudo-label collection from self-training. 2) Multiple factors jointly contributing to how the object is perceived in the unseen target domain. In this work, we propose a grouping-exploration strategy framework, Group Explorer Domain Adaptation ($\textbf{GroupEXP-DA}$), to addresses those two issues. Specifically, our grouping divides the available label sets into multiple clusters and ensures all of them have equal learning attention with the group-equivariant spatial feature, avoiding dominant types of objects causing imbalance problems. Moreover, grouping learns to divide objects by considering inherent factors in a data-driven manner, without considering each factor separately as existing works. On top of the group-equivariant spatial feature that selectively detects objects similar to the input group, we additionally introduce an explorative group update strategy that reduces the false negative detection in the target domain, further reducing the inter-domain gap. During inference, only the learned group features are necessary for making the group-equivariant spatial feature, placing our method as a simple add-on that can be applicable to most existing detectors. We show how each module contributes to substantially bridging the inter-domain gaps compared to existing works across large urban outdoor datasets such as NuScenes, Waymo, and KITTI. Sang-Yun Shin, Madhu Vankadari, Ta Ying Cheng, Qian Xie 0001, Andrew Markham, Agathoniki Trigoni |
NeurIPS | 4 |
| 2024 | Beyond Fusion: Modality Hallucination-based Multispectral Fusion for Pedestrian DetectionabstractPedestrian detection is a fundamental task for many downstream applications. Visible and thermal images, as the two most important data types, are usually used to detect pedestrians under various environmental conditions. Many state-of-the-art works have been proposed to use two-stream (i.e., two-branch) architectures to combine visible and thermal information to improve detection performance. However, conventional visible-thermal fusion-based methods have no ability to obtain useful information from the visible branch under poor visibility conditions. The visible branch could even sometimes bring noise into the combined features. In this paper, we present a novel thermal and visible fusion architecture for pedestrian detection. Instead of simply using two branches to separately extract thermal and visible features and then fusing them, we introduce a hallucination branch to learn the mapping from the thermal to the visible domain, forming a novel three-branch feature extraction module. We then adaptively fuse feature maps from all three branches (i.e., thermal, visible, and hallucination). With this new integrated hallucination branch, our network can still get relatively good visible feature maps under challenging low-visibility conditions, thus boosting the overall detection performance. Finally, we experimentally demonstrate the superiority of the proposed architecture over conventional fusion methods. Qian Xie 0001, Ta Ying Cheng, Jia-Xing Zhong, Kaichen Zhou, Andrew Markham, Agathoniki Trigoni |
WACV | 2 |
| 2024 | Illumination-Aware Hallucination-Based Domain Adaptation for Thermal Pedestrian DetectionabstractThermal imagery is emerging as a viable candidate for 24-7, all-weather pedestrian detection owning to thermal sensors’ robust performance for pedestrian detection under different weather and illumination conditions. Despite the promising results obtained from combining visible (RGB) and thermal cameras in multi-spectral fusion techniques, the complex synchronization requirements, including alignment and calibration of sensors, impede their deployment in real-world scenarios. In this paper, we introduce a novel approach for domain adaptation to enhance the performance of pedestrian detection based solely on thermal images. Our proposed approach involves several stages. Firstly, we use both thermal and visible images as input during the training phase. Secondly, we leverage a thermal-to-visible hallucination network to generate feature maps that are similar to those generated by the visible branch. Finally, we design a transformer-based multi-modal fusion module to integrate the hallucinated visible and thermal information more effectively. The thermal-to-visible hallucination network acts as domain adaptation, allowing us to obtain pseudo-visual and thermal features using solely thermal input. Based on the experimental results, it is observed the mean average precision (mAP) increases by 4.72% and the miss rate decreases by 7.56% on the KAIST dataset when compared to the baseline model. Qian Xie 0001, Ta Ying Cheng, Zhuangzhuang Dai, Vu H. Tran, Agathoniki Trigoni, Andrew Markham |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | mmPoint: Dense Human Point Cloud Generation from mmWave
Qian Xie 0001, Qianyi Deng, Ta Ying Cheng, Peijun Zhao, Amir Patel, Agathoniki Trigoni, Andrew Markham |
BMVC | 3 |
| 2023 | 3DMiner: Discovering Shapes from Large-Scale Unannotated Image DatasetsabstractWe present 3DMiner – a pipeline for mining 3D shapes from challenging large-scale unannotated image datasets. Unlike other unsupervised 3D reconstruction methods, we assume that, within a large-enough dataset, there must exist images of objects with similar shapes but varying backgrounds, textures, and viewpoints. Our approach leverages the recent advances in learning self-supervised image representations to cluster images with geometrically similar shapes and find common image correspondences between them. We then exploit these correspondences to obtain rough camera estimates as initialization for bundle-adjustment. Finally, for every image cluster, we apply a progressive bundle-adjusting reconstruction method to learn a neural occupancy field representing the underlying shape. We show that this procedure is robust to several types of errors introduced in previous steps (e.g., wrong camera poses, images containing dissimilar shapes, etc.), allowing us to obtain shape and pose annotations for images in-the-wild. When using images from Pix3D chairs, our method is capable of producing significantly better results than state-of-the-art unsupervised 3D reconstruction techniques, both quantitatively and qualitatively. Furthermore, we show how 3DMiner can be applied to in-the-wild data by reconstructing shapes present in images from the LAION-5B dataset. Project Page: https://ttchengab.github.io/3dminerOfficial. Ta Ying Cheng, Matheus Gadelha, Sören Pirk, Thibault Groueix, Radomír Mech, Andrew Markham, Agathoniki Trigoni |
ICCV | 1 |
| 2023 | Multi-body SE(3) Equivariance for Unsupervised Rigid Segmentation and Motion EstimationabstractA truly generalizable approach to rigid segmentation and motion estimation is fundamental to 3D understanding of articulated objects and moving scenes. In view of the closely intertwined relationship between segmentation and motion estimates, we present an SE(3) equivariant architecture and a training strategy to tackle this task in an unsupervised manner. Our architecture is composed of two interconnected, lightweight heads. These heads predict segmentation masks using point-level invariant features and estimate motion from SE(3) equivariant features, all without the need for category information. Our training strategy is unified and can be implemented online, which jointly optimizes the predicted segmentation and motion by leveraging the interrelationships among scene flow, segmentation mask, and rigid transformations. We conduct experiments on four datasets to demonstrate the superiority of our method. The results show that our method excels in both model performance and computational efficiency, with only 0.25M parameters and 0.92G FLOPs. To the best of our knowledge, this is the first work designed for category-agnostic part-level SE(3) equivariance in dynamic point clouds. Jia-Xing Zhong, Ta Ying Cheng, Kai Lu 0003, Kaichen Zhou, Andrew Markham, Agathoniki Trigoni |
NeurIPS | 2 |
| 2023 | Efficient 3D Feature Learning for Real-Time AwarenessabstractThis extended abstract discusses the current methods and work progress on sampling large-scale point cloud datasets with semantics and reconstructing 3D objects from sparse inputs. In particular, we describe a proposed meta sampling strategy to quickly adapt sampling to multiple tasks and potential methods to improve multi-modal reconstruction. These methods could benefit immensely in creating in-depth situational awareness for challenging missions and rescues. Ta Ying Cheng |
SMARTCOMP | 1 |
| 2022 | Pose Adaptive Dual Mixup for Few-Shot Single-View 3D ReconstructionabstractWe present a pose adaptive few-shot learning procedure and a two-stage data interpolation regularization, termed Pose Adaptive Dual Mixup (PADMix), for single-image 3D reconstruction. While augmentations via interpolating feature-label pairs are effective in classification tasks, they fall short in shape predictions potentially due to inconsistencies between interpolated products of two images and volumes when rendering viewpoints are unknown. PADMix targets this issue with two sets of mixup procedures performed sequentially. We first perform an input mixup which, combined with a pose adaptive learning procedure, is helpful in learning 2D feature extraction and pose adaptive latent encoding. The stagewise training allows us to build upon the pose invariant representations to perform a follow-up latent mixup under one-to-one correspondences between features and ground-truth volumes. PADMix significantly outperforms previous literature on few-shot settings over the ShapeNet dataset and sets new benchmarks on the more challenging real-world Pix3D dataset. Ta Ying Cheng, Hsuan-Ru Yang, Agathoniki Trigoni, Hwann-Tzong Chen, Tyng-Luh Liu |
AAAI | 1 |
| 2022 | Meta-sampler: Almost-Universal yet Task-Oriented Sampling for Point Clouds
Ta Ying Cheng, Qingyong Hu, Qian Xie 0001, Agathoniki Trigoni, Andrew Markham |
ECCV (2) | 1 |
| 2020 | ARchitect: Building Interactive Virtual Experiences from Physical Affordances by Bringing Human-in-the-LoopabstractAutomatic generation of Virtual Reality (VR) worlds which adapt to physical environments have been proposed to enable safe walking in VR. However, such techniques mainly focus on the avoidance of physical objects as obstacles and overlook their interaction affordances as passive haptics. Current VR experiences involving interaction with physical objects in surroundings still require verbal instruction from an assisting partner. We present ARchitect, a proof-of-concept prototype that allows flexible customization of a VR experience with human-in-the-loop. ARchitect brings in an assistant to map physical objects to virtual proxies of matching affordances using Augmented Reality (AR). In a within-subjects study (9 user pairs) comparing ARchitect to a baseline condition, assistants and players experienced decreased workload and players showed increased VR presence and trust in the assistant. Finally, we defined design guidelines of ARchitect for future designers and implemented three demonstrative experiences. David Chuan-En Lin, Ta Ying Cheng, Xiaojuan Ma |
CHI | 2 |
| 2020 | SeqDynamics: Visual Analytics for Evaluating Online Problem-solving DynamicsabstractAbstract Problem‐solving dynamics refers to the process of solving a series of problems over time, from which a student's cognitive skills and non‐cognitive traits and behaviors can be inferred. For example, we can derive a student's learning curve (an indicator of cognitive skill) from the changes in the difficulty level of problems solved, or derive a student's self‐regulation patterns (an example of non‐cognitive traits and behaviors) based on the problem‐solving frequency over time. Few studies provide an integrated overview of both aspects by unfolding the problem‐solving process. In this paper, we present a visual analytics system named SeqDynamics that evaluates students ‘problem‐solving dynamics from both cognitive and non‐cognitive perspectives. The system visualizes the chronological sequence of learners’ problem‐solving behavior through a set of novel visual designs and coordinated contextual views, enabling users to compare and evaluate problem‐solving dynamics on multiple scales. We present three scenarios to demonstrate the usefulness of SeqDynamics on a real‐world dataset which consists of thousands of problem‐solving traces. We also conduct five expert interviews to show that SeqDynamics enhances domain experts’ understanding of learning behavior sequences and assists them in completing evaluation tasks efficiently. Meng Xia 0002, David Chuan-En Lin, Ta Ying Cheng, Huamin Qu, Xiaojuan Ma |
Comput. Graph. Forum | 4 |