Hang Chu

dblp:64/10446 · DBLP profile ↗
← Back
18ranked-venue papers
9as first author
6since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 7 first-author · 5 since 2021Artificial intelligence and machine learning · 14 · 8 first-author · 5 since 2021Systems, architecture and hardware · 2 · 1 first-author
YearPublicationVenuePosition
2022 CLIP-Forge: Towards Zero-Shot Text-to-Shape Generation
abstract
Generating shapes using natural language can enable new ways of imagining and creating the things around us. While significant recent progress has been made in text-to-image generation, text-to-shape generation remains a challenging problem due to the unavailability of paired text and shape data at a large scale. We present a simple yet effective method for zero-shot text-to-shape gener-ation that circumvents such data scarcity. Our proposed method, named CLIP-Forge, is based on a two-stage training process, which only depends on an unlabelled shape dataset and a pre-trained image-text network such as CLIP. Our method has the benefits of avoiding expensive inference time optimization, as well as the ability to generate multiple shapes for a given text. We not only demonstrate promising zero-shot generalization of the CLIP-Forge model qualitatively and quantitatively, but also provide extensive compar-ative evaluations to better understand its behavior.
Aditya Sanghi, Hang Chu, Joseph G. Lambourne, Chin-Yi Cheng, Marco Fumero, Kamal Rahimi Malekshan
CVPR2
2022 JoinABLe: Learning Bottom-up Assembly of Parametric CAD Joints
abstract
Physical products are often complex assemblies combining a multitude of 3D parts modeled in computer-aided design (CAD) software. CAD designers build up these assemblies by aligning individual parts to one another using constraints called joints. In this paper we introduce JoinABLe, a learning-based method that assembles parts together to form joints. JoinABLe uses the weak supervision available in standard parametric CAD files without the help of object class labels or human guidance. Our results show that by making network predictions over a graph representation of solid models we can outperform multiple baseline methods with an accuracy (79.53%) that approaches human performance (80%). Finally, to support future research we release the Fusion 360 Gallery assembly dataset, containing assemblies with rich information on joints, contact surfaces, holes, and the underlying assembly graph structure.
Karl D. D. Willis, Pradeep Kumar Jayaraman, Hang Chu, Yunsheng Tian, Yifei Li 0002, Daniele Grandi, Aditya Sanghi, Joseph G. Lambourne, Armando Solar-Lezama, Wojciech Matusik
CVPR3
2022 SimCURL: Simple Contrastive User Representation Learning from Command Sequences
abstract
User modeling is crucial to understanding user behavior and essential for improving user experience and personalized recommendations. When users interact with software, vast amounts of command sequences are generated through logging and analytics systems. These command sequences contain clues to the users’ goals and intents. However, these data modalities are highly unstructured and unlabeled, making it difficult for standard predictive systems to learn from. We propose SimCURL, a simple yet effective contrastive self-supervised deep learning framework that learns user representation from unlabeled command sequences. Our method introduces a user-session network architecture, as well as session dropout as a novel way of data augmentation. We train and evaluate our method on a real-world command sequence dataset of more than half a billion commands. Our method shows significant improvement over existing methods when the learned representation is transferred to downstream tasks such as experience and expertise classification.
Hang Chu, Amir Khasahmadi, Karl D. D. Willis, Fraser Anderson, Yaoli Mao, Justin Matejka, Jo Vermeulen
ICMLA1
2021 House-GAN++: Generative Adversarial Layout Refinement Network towards Intelligent Computational Agent for Professional Architects
abstract
This paper proposes a generative adversarial layout refinement network for automated floorplan generation. Our architecture is an integration of a graph-constrained relational GAN and a conditional GAN, where a previously generated layout becomes the next input constraint, enabling iterative refinement. A surprising discovery of our research is that a simple non-iterative training process, dubbed component-wise GT-conditioning, is effective in learning such a generator. The iterative generator further allows us to improve a metric of choice via meta-optimization techniques by controlling when to pass which input constraints during iterative refinement. Our qualitative and quantitative evaluation based on the three standard metrics demonstrate that the proposed system makes significant improvements over the current state-of-the-art, even competitive against the ground-truth floorplans, designed by professional architects. Code, model, and data are available at https://ennauata.github.io/houseganpp/page.html.
Nelson Nauata, Sepidehsadat Hosseini, Kai-Hung Chang, Hang Chu, Chin-Yi Cheng, Yasutaka Furukawa
CVPR4
2021 LSD-StructureNet: Modeling Levels of Structural Detail in 3D Part Hierarchies
abstract
Generative models for 3D shapes represented by hierarchies of parts can generate realistic and diverse sets of outputs. However, existing models suffer from the key practical limitation of modelling shapes holistically and thus cannot perform conditional sampling, i.e. they are not able to generate variants on individual parts of generated shapes without modifying the rest of the shape. This is limiting for applications such as 3D CAD design that involve adjusting created shapes at multiple levels of detail. To address this, we introduce LSD-StructureNet, an augmentation to the StructureNet architecture that enables re-generation of parts situated at arbitrary positions in the hierarchies of its outputs. We achieve this by learning individual, probabilistic conditional decoders for each hierarchy depth. We evaluate LSD-StructureNet on the PartNet dataset, the largest dataset of 3D shapes represented by hierarchies of parts. Our results show that contrarily to existing methods, LSD-StructureNet can perform conditional sampling without impacting inference speed or the realism and diversity of its outputs.
Dominic Roberts, Ara Danielyan, Hang Chu, Mani Golparvar Fard, David A. Forsyth
ICCV3
2021 Fusion 360 gallery: a dataset and environment for programmatic CAD construction from human design sequences
abstract
Parametric computer-aided design (CAD) is a standard paradigm used to design manufactured objects, where a 3D shape is represented as a program supported by the CAD software. Despite the pervasiveness of parametric CAD and a growing interest from the research community, currently there does not exist a dataset of realistic CAD models in a concise programmatic form. In this paper we present the Fusion 360 Gallery , consisting of a simple language with just the sketch and extrude modeling operations, and a dataset of 8,625 human design sequences expressed in this language. We also present an interactive environment called the Fusion 360 Gym , which exposes the sequential construction of a CAD program as a Markov decision process, making it amendable to machine learning approaches. As a use case for our dataset and environment, we define the CAD reconstruction task of recovering a CAD program from a target geometry. We report results of applying state-of-the-art methods of program synthesis with neurally guided search on this task.
Karl D. D. Willis, Yewen Pu, Jieliang Luo, Hang Chu, Tao Du 0001, Joseph G. Lambourne, Armando Solar-Lezama, Wojciech Matusik
ACM Trans. Graph.4
2020 Expressive Telepresence via Modular Codec Avatars
Hang Chu, Shugao Ma, Fernando De la Torre, Sanja Fidler, Yaser Sheikh
ECCV (12)1
2019 Neural Turtle Graphics for Modeling City Road Layouts
abstract
We propose Neural Turtle Graphics (NTG), a novel generative model for spatial graphs, and demonstrate its applications in modeling city road layouts. Specifically, we represent the road layout using a graph where nodes in the graph represent control points and edges in the graph represents road segments. NTG is a sequential generative model parameterized by a neural network. It iteratively generates a new node and an edge connecting to an existing node conditioned on the current graph. We train NTG on Open Street Map data and show it outperforms existing approaches using a set of diverse performance metrics. Moreover, our method allows users to control styles of generated road layouts mimicking existing cities as well as to sketch a part of the city road layout to be synthesized. In addition to synthesis, the proposed NTG finds uses in an analytical task of aerial road parsing. Experimental results show that it achieves state-of-the-art performance on the SpaceNet dataset.
Hang Chu, Daiqing Li, David Acuna, Amlan Kar, Maria Shugrina, Xinkai Wei, Ming-Yu Liu 0001, Antonio Torralba 0001, Sanja Fidler
ICCV1
2018 A Face-to-Face Neural Conversation Model
abstract
Neural networks have recently become good at engaging in dialog. However, current approaches are based solely on verbal text, lacking the richness of a real face-to-face conversation. We propose a neural conversation model that aims to read and generate facial gestures alongside with text. This allows our model to adapt its response based on the "mood" of the conversation. In particular, we introduce an RNN encoder-decoder that exploits the movement of facial muscles, as well as the verbal conversation. The decoder consists of two layers, where the lower layer aims at generating the verbal response and coarse facial expressions, while the second layer fills in the subtle gestures, making the generated output more smooth and natural. We train our neural network by having it "watch" 250 movies. We showcase our joint face-text model in generating more natural conversations through automatic metrics and a human study. We demonstrate an example application with a face-to-face chatting avatar.
Hang Chu, Daiqing Li, Sanja Fidler
CVPR1
2018 SurfConv: Bridging 3D and 2D Convolution for RGBD Images
abstract
The last few years have seen approaches trying to combine the increasing popularity of depth sensors and the success of the convolutional neural networks. Using depth as additional channel alongside the RGB input has the scale variance problem present in image convolution based approaches. On the other hand, 3D convolution wastes a large amount of memory on mostly unoccupied 3D space, which consists of only the surface visible to the sensor. Instead, we propose SurfConv, which "slides" compact 2D filters along the visible 3D surface. SurfConv is formulated as a simple depth-aware multi-scale 2D convolution, through a new Data-Driven Depth Discretization (D4) scheme. We demonstrate the effectiveness of our method on indoor and outdoor 3D semantic segmentation datasets. Our method achieves state-of-the-art performance while using less than 30% parameters used by the 3D convolution based approaches.
Hang Chu, Wei-Chiu Ma, Kaustav Kundu, Raquel Urtasun, Sanja Fidler
CVPR1
2018 Single Image Intrinsic Decomposition Without a Single Intrinsic Image
Wei-Chiu Ma, Hang Chu, Bolei Zhou, Raquel Urtasun, Antonio Torralba 0001
ECCV (14)2
2017 TorontoCity: Seeing the World with a Million Eyes
abstract
In this paper we introduce the TorontoCity benchmark, which covers the full greater Toronto area (GTA) with 712.5km2 of land, 8439km of road and around 400, 000 buildings. Our benchmark provides different perspectives of the world captured from airplanes, drones and cars driving around the city. Manually labeling such a large scale dataset is infeasible. Instead, we propose to utilize different sources of high-precision maps to create our ground truth. Towards this goal, we develop algorithms that allow us to align all data sources with the maps while requiring minimal human supervision. We have designed a wide variety of tasks including building height estimation (reconstruction), road centerline and curb extraction, building instance segmentation, building contour extraction (reorganization), semantic labeling and scene type classification (recognition). Our pilot study shows that most of these tasks are still difficult for modern convolutional neural networks.
Shenlong Wang, Min Bai, Gellért Máttyus, Hang Chu, Wenjie Luo 0002, Bin Yang 0021, Justin Liang 0001, Joel Cheverie, Sanja Fidler, Raquel Urtasun
ICCV4
2016 HouseCraft: Building Houses from Rental Ads and Street Views
Hang Chu, Shenlong Wang, Raquel Urtasun, Sanja Fidler
ECCV (6)1
2015 You are Here: Mimicking the Human Thinking Process in Reading Floor-Plans
abstract
A human can easily find his or her way in an unfamiliar building, by walking around and reading the floor-plan. We try to mimic and automate this human thinking process. More precisely, we introduce a new and useful task of locating an user in the floor-plan, by using only a camera and a floor-plan without any other prior information. We address the problem with a novel matching-localization algorithm that is inspired by human logic. We demonstrate through experiments that our method outperforms state-of-the-art floor-plan-based localization methods by a large margin, while also being highly efficient for real-time applications.
Hang Chu, Dong Ki Kim, Tsuhan Chen
ICCV1
2015 Consistent ground-plane mapping: A case study utilizing low-cost sensor measurements and a satellite image
abstract
Vision-aided localization systems are often utilized in urban settings to take advantage of structured environment, high availability of unique visual features, as well as complimenting aiding measurements from Global Navigation Satellite System (GNSS). In this paper, we present a case study for roadway texture mapping that combines low-cost sensor measurements that are already available on many production vehicles (e.g. single frequency GPS, wheel odometry, and a forward looking camera) together with a satellite image. The aim of the method presented here is to obtain high resolution texture of the ground plane that is consistent with the low-resolution satellite image through an optimization process that estimates the smooth vehicle trajectory using Maximum-a-Posteriori (MAP). The main benefit of this system comes from the facts that: (1) it utilizes only low-cost sensors and information that are readily available, (2) it can be easily embedded into existing maps. Data and analysis of a drive captured around a block is used in this study.
Hang Chu, Anh Vu
ICRA1
2013 A new Local-Main-Gradient-Orientation HOG and contour differences based algorithm for object classification
abstract
This paper presents a new algorithm to better classify objects in videos. In our case, the objects are cars, vans, and people on the roads. First, in order to extract the moving objects more precisely, we have proposed a method for foreground extraction based on the contour differences between the video frame and the background image. Second, after we got the integrated moving object, we have proposed a new algorithm to extract better features from the object. The new algorithm is based on two extended Histogram of Oriented Gradient (HOG) descriptor. We have improved HOG in two aspects: (a) selecting the gradient information from the moving objects and discarding the background gradient; (b) weighting every bin of gradient orientation histogram according to their significance within predefined area, in order to emphasize the important gradient information. We obtained Contour-Difference HOG (CD-HOG) from the first extension and Local-Main-Gradient-Orientation HOG (LMGO-HOG) from the second extended HOG. These extensions can cope with the cluttered background and make the features more distinguishable. Each of the extended HOG descriptors can produce a satisfying performance separately and an even better one if they are applied in cascade. From extensive evaluations, we showed the wonderful performance of our algorithm, and the accuracy rate of 94.04% can be achieved in some cases.
Xiaoqiong Su, Weiyao Lin, Xiaozhen Zheng, Xintong Han, Hang Chu, Xiaoyun Zhang 0001
ISCAS5
2013 A Heat-Map-Based Algorithm for Recognizing Group Activities in Videos
abstract
In this paper, a new heat-map-based algorithm is proposed for group activity recognition. The proposed algorithm first models human trajectories as series of heat sources and then applies a thermal diffusion process to create a heat map (HM) for representing the group activities. Based on this HM, a new key-point-based (KPB) method is used for handling the alignments among HMs with different scales and rotations. A surface-fitting (SF) method is also proposed for recognizing group activities. Our proposed HM feature can efficiently embed the temporal motion information of the group activities while the proposed KPB and SF methods can effectively utilize the characteristics of the HM for activity recognition. Section IV demonstrates the effectiveness of our proposed algorithms.
Weiyao Lin, Hang Chu, Jianxin Wu 0001, Bin Sheng 0001, Zhenzhong Chen 0001
IEEE Trans. Circuits Syst. Video Technol.2
2012 A new heat-map-based algorithm for human group activity recognition
abstract
In this paper, a new heat-map-based (HMB) algorithm is proposed for human group activity recognition. The proposed algorithm first models people trajectories as series of "heat sources" and then applies a thermal diffusion process to create a heat map (HM) for representing the group activities. Based on this heat map, a new surface-fitting (SF) method is also proposed for recognizing human group activities. Our proposed HM feature can efficiently keep the temporal motion information of the group activities while the proposed SF method can effectively catch the characteristics of the heat map for activity recognition. Experimental results demonstrate the effectiveness of our proposed algorithm.
Hang Chu, Weiyao Lin, Jianxin Wu 0001, Xingtong Zhou, Yuanzhe Chen, Hongxiang Li 0001
ACM Multimedia1