Cheng-Chun Hsu

dblp:228/1377 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Robot manipulation · 49% 3D vision · 18% Segmentation and scene understanding · 15%
Databases, data mining, and information retrieval
2 papers
Recommender systems · 77% Web and social media mining · 23%
Computer graphics and multimedia
1 paper
Multimedia analysis and retrieval · 100%

Topics — the 19 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Robot manipulation › grasping
articulated object manipulation
1.222023
Ditto in the House: Building Articulation Models of Indoor Scenes through Interactive Perception · ICRA 2023
Ditto: Building Digital Twins of Articulated Objects from Interaction · CVPR 2022
Robotics › Robot manipulation › learning from demonstration
imitation learning for manipulation
0.912025
SPOT: SE(3) Pose Trajectory Diffusion for Object-Centric Manipulation · ICRA 2025
Robotics › Robot manipulation › object manipulation
object-centric manipulation
0.912025
SPOT: SE(3) Pose Trajectory Diffusion for Object-Centric Manipulation · ICRA 2025
Computer vision › Segmentation and scene understanding
scene understanding
0.712023
Ditto in the House: Building Articulation Models of Indoor Scenes through Interactive Perception · ICRA 2023
Computer vision › 3D vision
3d reconstruction
0.612022
Ditto: Building Digital Twins of Articulated Objects from Interaction · CVPR 2022
Computer vision › 3D vision
3d scene understanding
0.612022
Ditto: Building Digital Twins of Articulated Objects from Interaction · CVPR 2022
Computer vision › 3D vision › 3d reconstruction › object reconstruction
articulated object reconstruction
0.612022
Ditto: Building Digital Twins of Articulated Objects from Interaction · CVPR 2022
Robotics › Robot manipulation
grasping
0.612022
Ditto: Building Digital Twins of Articulated Objects from Interaction · CVPR 2022
Robotics › Robot manipulation › robot sensing › perception for manipulation
interactive perception
0.612022
Ditto: Building Digital Twins of Articulated Objects from Interaction · CVPR 2022
Multimedia analysis and retrieval › multimodal learning
multimodal representation learning
0.512021
Dress With Style: Learning Style From Joint Deep Embedding of Clothing Styles and Body Shapes · IEEE Trans. Multim. 2021
Computer vision › Image recognition and object detection › object detection
domain adaptive object detection
0.412020
Every Pixel Matters: Center-Aware Feature Alignment for Domain Adaptive Object Detector · ECCV (9) 2020
Machine learning › Representation and self-supervised learning › representation matching
feature alignment
0.412020
Every Pixel Matters: Center-Aware Feature Alignment for Domain Adaptive Object Detector · ECCV (9) 2020
Computer vision › Image recognition and object detection
object detection
0.412020
Every Pixel Matters: Center-Aware Feature Alignment for Domain Adaptive Object Detector · ECCV (9) 2020
Computer vision › Segmentation and scene understanding
instance segmentation
0.412019
Weakly Supervised Instance Segmentation using the Bounding Box Tightness Prior · NeurIPS 2019
Machine learning › Learning paradigms
multiple instance learning
0.412019
Weakly Supervised Instance Segmentation using the Bounding Box Tightness Prior · NeurIPS 2019
Computer vision › Segmentation and scene understanding › instance segmentation
weakly supervised instance segmentation
0.412019
Weakly Supervised Instance Segmentation using the Bounding Box Tightness Prior · NeurIPS 2019
Recommender systems
fashion recommendation
0.312018
What Dress Fits Me Best?: Fashion Recommendation on the Clothing Style for Personal Body Shape · ACM Multimedia 2018
Robotics › Robot manipulation
diffusion policy
0.312025
SPOT: SE(3) Pose Trajectory Diffusion for Object-Centric Manipulation · ICRA 2025
Robotics › Robot manipulation › object perception
affordance detection
0.212023
Ditto in the House: Building Articulation Models of Indoor Scenes through Interactive Perception · ICRA 2023

Methods — techniques the papers use, named apart from their topics

graph propagation · 1.0deep multimodal representation learning · 1.0diffusion policy · 0.9SE(3) pose trajectory representation · 0.9interactive perception · 0.7articulation inference · 0.7affordance prediction · 0.7physical simulation · 0.6implicit neural representation · 0.6center-aware feature alignment · 0.4multiple instance learning · 0.4deep neural network · 0.4
YearPublicationVenuePosition
2025 SPOT: SE(3) Pose Trajectory Diffusion for Object-Centric Manipulation
abstract
We introduce SPOT, an object-centric imitation learning framework. The key idea is to capture each task by an object-centric representation, specifically the SE(3) object pose trajectory relative to the target. This approach decouples embodiment actions from sensory inputs, facilitating learning from various demonstration types, including both action-based and action-less human hand demonstrations, as well as crossembodiment generalization. Additionally, object pose trajectories inherently capture planning constraints from demonstrations without the need for manually-crafted rules. To guide the robot in executing the task, the object trajectory is used to condition a diffusion policy. We systematically evaluate our method on simulation and real-world tasks. In real-world evaluation, using only eight demonstrations shot on an iPhone, our approach completed all tasks while fully complying with task constraints. Project page: https://nvlabs.github.io/object_centric_diffusion
Cheng-Chun Hsu, Bowen Wen, Jie Xu 0028, Yashraj S. Narang, Xiaolong Wang 0004, Yuke Zhu, Joydeep Biswas, Stanley T. Birchfield
ICRA1
2023 Ditto in the House: Building Articulation Models of Indoor Scenes through Interactive Perception
abstract
Virtualizing the physical world into virtual models has been a critical technique for robot navigation and planning in the real world. To foster manipulation with articulated objects in everyday life, this work explores building articulation models of indoor scenes through a robot's purposeful inter-actions in these scenes. Prior work on articulation reasoning primarily focuses on siloed objects of limited categories. To extend to room-scale environments, the robot has to efficiently and effectively explore a large-scale 3D space, locate articulated objects, and infer their articulations. We introduce an interactive perception approach to this task. Our approach, named Ditto in the House, discovers possible articulated objects through affordance prediction, interacts with these objects to produce articulated motions, and infers the articulation properties from the visual observations before and after each interaction. It tightly couples affordance prediction and articulation inference to improve both tasks. We demonstrate the effectiveness of our approach in both simulation and real-world scenes. Code and additional results are available at https://ut-austin-rpl.github.io/HouseDitto/
Cheng-Chun Hsu, Zhenyu Jiang 0002, Yuke Zhu
ICRA1
2022 Ditto: Building Digital Twins of Articulated Objects from Interaction
abstract
Digitizing physical objects into the virtual world has the potential to unlock new research and applications in embodied AI and mixed reality. This work focuses on recreating interactive digital twins of real-world articulated objects, which can be directly imported into virtual environments. We introduce Ditto to learn articulation model estimation and 3D geometry reconstruction of an articulated object through interactive perception. Given a pair of visual observations of an articulated object before and after interaction, Ditto reconstructs part-level geometry and estimates the articulation model of the object. We employ implicit neural representations for joint geometry and articulation modeling. Our experiments show that Ditto effectively builds digital twins of articulated objects in a category-agnostic way. We also apply Ditto to real-world objects and deploy the recreated digital twins in physical simulation. Code and additional results are available at https://ut-austin-rpl.github.io/Ditto/
Zhenyu Jiang 0002, Cheng-Chun Hsu, Yuke Zhu
CVPR2
2021 Dress With Style: Learning Style From Joint Deep Embedding of Clothing Styles and Body Shapes
abstract
Body shape is about proportion, and fashion style is all about dressing those proportions to look their very best. Figuring out the styles to suit a body shape can be a daunting task for many people. It is, therefore, essential to develop a framework for learning the compatibility of body shapes and clothing styles. Though fashion designers and fashion stylists have analyzed the correlation between human body shapes and fashion styles for a long time, this issue did not receive much attention in multimedia science. In this paper, we present a novel style recommender, on the basis of the user's body attributes. The rich amount of fashion styling knowledge from social big data is exploited for this purpose. We first construct a joint embedding of clothing styles and human body measurements with deep multimodal representation learning on a reference dataset that has been sorted to meet the fashion rules. We then discover the relevant semantic features by propagation and selection in clothing style and body shape graphs. Experiments demonstrate the effectiveness of the proposed framework when compared with several baseline methods.
Shintami Chusnul Hidayati, Ting Wei Goh, Ji-Sheng Gary Chan, Cheng-Chun Hsu, John See, Lai-Kuan Wong, Kai-Lung Hua, Yu Tsao 0001, Wen-Huang Cheng
IEEE Trans. Multim.4
2020 Every Pixel Matters: Center-Aware Feature Alignment for Domain Adaptive Object Detector
Cheng-Chun Hsu, Yi-Hsuan Tsai, Yen-Yu Lin, Ming-Hsuan Yang 0001
ECCV (9)1
2019 Weakly Supervised Instance Segmentation using the Bounding Box Tightness Prior
abstract
This paper presents a weakly supervised instance segmentation method that consumes training data with tight bounding box annotations. The major difficulty lies in the uncertain figure-ground separation within each bounding box since there is no supervisory signal about it. We address the difficulty by formulating the problem as a multiple instance learning (MIL) task, and generate positive and negative bags based on the sweeping lines of each bounding box. The proposed deep model integrates MIL into a fully supervised instance segmentation network, and can be derived by the objective consisting of two terms, i.e., the unary term and the pairwise term. The former estimates the foreground and background areas of each bounding box while the latter maintains the unity of the estimated object masks. The experimental results show that our method performs favorably against existing weakly supervised methods and even surpasses some fully supervised methods for instance segmentation on the PASCAL VOC dataset.
Cheng-Chun Hsu, Kuang-Jui Hsu, Chung-Chi Tsai, Yen-Yu Lin, Yung-Yu Chuang
NeurIPS1
2018 What Dress Fits Me Best?: Fashion Recommendation on the Clothing Style for Personal Body Shape
abstract
Clothing is an integral part of life. Also, it is always an uneasy task for people to make decisions on what to wear. An essential style tip is to dress for the body shape, i.e., knowing one's own body shape (e.g., hourglass, rectangle, round and inverted triangle) and selecting the types of clothes that will accentuate the body's good features. In the literature, although various fashion recommendation systems for clothing items have been developed, none of them had explicitly taken the user's basic body shape into consideration. In this paper, therefore, we proposed a first framework for learning the compatibility of clothing styles and body shapes from social big data, with the goal to recommend a user about what to wear better in relation to his/her essential body attributes. The experimental results demonstrate the superiority of our proposed approach, leading to a new aspect for research into fashion recommendation.
Shintami Chusnul Hidayati, Cheng-Chun Hsu, Yu-Ting Chang, Kai-Lung Hua, Jianlong Fu, Wen-Huang Cheng
ACM Multimedia2