EDBT 2026 Demo / reviewers in the wild / expert
Tai-Jiang Mu
dblp:146/4849 · also Taijiang Mu
· DBLP profile ↗
64ranked-venue papers
4as first author
49since 2021 · last 2025
0000-0002-9197-346XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 55 · 4 first-author · 43 since 2021Artificial intelligence and machine learning · 16 · 13 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | High-Accuracy Fractured Object Reassembly Under Arbitrary Poses
Qun-Ce Xu, Yan-Pei Cao 0001, Weihao Cheng 0002, Tai-Jiang Mu, Ying Shan, Yongliang Yang 0002, Shi-Min Hu 0001 |
CVM (2) | 4 |
| 2025 | RGE-GS: Reward-Guided Expansive Driving Scene Reconstruction via Diffusion Priors
Sicong Du, Jiarun Liu, Haoxiang Chen 0004, Tai-Jiang Mu, Sheng Yang 0007 |
ICCV | 5 |
| 2025 | SDLKF: Signed Distance Linear Kernel Function for surface reconstruction
Haoxiang Chen 0004, Xiao-Lei Li, Tai-Jiang Mu, Qun-Ce Xu, Shi-Min Hu 0001 |
Comput. Graph. | 3 |
| 2025 | RS-SpecSDF: Reflection-supervised surface reconstruction and material estimation for specular indoor scenesabstractNeural Radiance Field (NeRF) has achieved impressive 3D reconstruction quality using implicit scene representations. However, planar specular reflections pose significant challenges in the 3D reconstruction task. It is a common practice to decompose the scene into physically real geometries and virtual images produced by the reflections. However, current methods struggle to resolve the ambiguities in the decomposition process , because they mostly rely on mirror masks as external cues. They also fail to acquire accurate surface materials, which is essential for downstream applications of the recovered geometries. In this paper, we present RS-SpecSDF, a novel framework for indoor scene surface reconstruction that can faithfully reconstruct specular reflectors while accurately decomposing the reflection from the scene geometries and recovering the accurate specular fraction and diffuse appearance of the surface without requiring mirror masks. Our key idea is to perform reflection ray-casting and use it as supervision for the decomposition of reflection and surface material. Our method is based on an observation that the virtual image seen by the camera ray should be consistent with the object that the ray hits after reflecting off the specular surface. To leverage this constraint, we propose the Reflection Consistency Loss and Reflection Certainty Loss to regularize the decomposition. Experiments conducted on both our newly-proposed synthetic dataset and a real-captured dataset demonstrate that our method achieves high-quality surface reconstruction and accurate material decomposition results without the need of mirror masks. Dong-Yu Chen, Haoxiang Chen 0004, Qun-Ce Xu, Tai-Jiang Mu |
Graph. Model. | 4 |
| 2025 | EasyAnim: 3D facial animation from in-the-wild videos for avatars with customized riggings
Haoxuan Song, Xiaohang Zhan, Tai-Jiang Mu |
Graph. Model. | 4 |
| 2025 | TerraCraft: City-scale generative procedural modeling with natural languagesabstractAutomated generation of large-scale 3D scenes presents a significant challenge due to the resource-intensive training and datasets required. This is in sharp contrast to the 2D counterparts that have become readily available due to their superior speed and quality. However, prior work in 3D procedural modeling has demonstrated promise in generating high-quality assets using the combination of algorithms and user-defined rules. To leverage the best of both 2D generative models and procedural modeling tools, we present TerraCraft, a novel framework for generating geometrically high-quality 3D city-scale scenes. By utilizing Large Language Models (LLMs), TerraCraft can generate city-scale 3D scenes from natural text descriptions. With its intuitive operation and powerful capabilities, TerraCraft enables users to easily create geometrically high-quality scenes readily for various applications, such as virtual reality and game design. We validate TerraCraft’s effectiveness through extensive experiments and user studies, showing its superior performance compared to existing baselines. Zhihao Yao 0004, Zi-Qi Lu, Hongyu Yan, Tai-Jiang Mu, Qun-Ce Xu |
Graph. Model. | 6 |
| 2025 | SN$^{2}$2eRF: A Framework for Neural Radiance Fields Given Sparse and Noisy PosesabstractNeural Radiance Fields (NeRFs) have shown impressive capabilities in synthesizing photorealistic novel views. However, their application to room-size scenes is limited by the requirement of several hundred views with accurate poses for training. To address this challenge, we propose SN$^{2}$2eRF, a framework which can reconstruct the neural radiance field with significantly fewer views and noisy poses by exploiting multiple priors. Our key insight is to leverage both multi-view and monocular priors to constrain the optimization of NeRF in the setting of sparse and noisy pose inputs. Specifically, we extract and match key points to constrain pose optimization and use Ray Transformer with a monocular depth estimator to provide dense depth prior for geometry optimization. Benefiting from these priors, our approach achieves state-of-the-art accuracy in novel view synthesis for indoor room scenarios. Haoxiang Chen 0004, Tai-Jiang Mu, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2025 | SLS4D: Sparse Latent Space for 4D Novel View SynthesisabstractNeural radiance fields (NeRF) have achieved great success in novel view synthesis and 3D representation for static scenarios. Existing dynamic NeRFs usually exploit a locally dense grid to fit the deformation fields; however, they fail to capture the global dynamics and concomitantly yield models of heavy parameters. We observe that the 4D space is inherently sparse. First, the deformation fields are sparse in spatial but dense in temporal due to the continuity of motion. Second, the radiance fields are only valid on the surface of the underlying scene, usually occupying a small fraction of the whole space. We thus represent the 4D scene using a learnable sparse latent space, a.k.a. SLS4D. Specifically, SLS4D first uses dense learnable time slot features to depict the temporal space, from which the deformation fields are fitted with linear multi-layer perceptions (MLP) to predict the displacement of a 3D position at any time. It then learns the spatial features of a 3D position using another sparse latent space. This is achieved by learning the adaptive weights of each latent feature with the attention mechanism. Extensive experiments demonstrate the effectiveness of our SLS4D: It achieves the best 4D novel view synthesis using only about 6% parameters of the most recent work. Qi-Yuan Feng, Haoxiang Chen 0004, Qun-Ce Xu, Tai-Jiang Mu |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | GP-Recon: Online Monocular Neural 3D Reconstruction With Geometric PriorabstractHigh-fidelity online 3D scene reconstruction from monocular videos continues to be challenging, especially for coherent and fine-grained geometry reconstruction. The previous learning-based online 3D reconstruction approaches with neural implicit representations have shown a promising ability for coherent scene reconstruction, but often fail to consistently reconstruct fine-grained geometric details during online reconstruction. This paper presents a new on-the-fly monocular 3D reconstruction approach, named GP-Recon, to perform high-fidelity online neural 3D reconstruction with fine-grained geometric details. We incorporate geometric prior (GP) into a scene's neural geometry learning to better capture its geometric details and, more importantly, propose an online volume rendering optimization to reconstruct and maintain geometric details during the online reconstruction task. The extensive comparisons with state-of-the-art approaches show that our GP-Recon consistently generates more accurate and complete reconstruction results with much better fine-grained details, both quantitatively and qualitatively. Zixin Zou, Shi-Sheng Huang, Yan-Pei Cao 0001, Tai-Jiang Mu, Ying Shan, Hongbo Fu 0001, Song-Hai Zhang |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | Semantic-Aware Transformation-Invariant RoI AlignabstractGreat progress has been made in learning-based object detection methods in the last decade. Two-stage detectors often have higher detection accuracy than one-stage detectors, due to the use of region of interest (RoI) feature extractors which extract transformation-invariant RoI features for different RoI proposals, making refinement of bounding boxes and prediction of object categories more robust and accurate. However, previous RoI feature extractors can only extract invariant features under limited transformations. In this paper, we propose a novel RoI feature extractor, termed Semantic RoI Align (SRA), which is capable of extracting invariant RoI features under a variety of transformations for two-stage detectors. Specifically, we propose a semantic attention module to adaptively determine different sampling areas by leveraging the global and local semantic relationship within the RoI. We also propose a Dynamic Feature Sampler which dynamically samples features based on the RoI aspect ratio to enhance the efficiency of SRA, and a new position embedding, i.e., Area Embedding, to provide more accurate position information for SRA through an improved sampling area representation. Experiments show that our model significantly outperforms baseline models with slight computational overhead. In addition, it shows excellent generalization ability and can be used to improve performance with various state-of-the-art backbones and detection methods. The code is available at https://github.com/cxjyxxme/SemanticRoIAlign. Guo-Ye Yang, George Kiyohiro Nakayama, Zi-Kai Xiao, Tai-Jiang Mu, Sharon X. Huang, Shi-Min Hu 0001 |
AAAI | 4 |
| 2024 | Programmable Motion Generation for Open-Set Motion Control TasksabstractCharacter animation in real-world scenarios necessitates a variety of constraints, such as trajectories, keyframes, interactions, etc. Existing methodologies typically treat single or a finite set of these constraint(s) as separate control tasks. These methods are often specialized, and the tasks they address are rarely extendable or customizable. We categorize these as solutions to the close-set motion control problem. In response to the complexity of practical motion control, we propose and attempt to solve the open-set motion control problem. This problem is characterized by an open and fully customizable set of motion control tasks. To address this, we introduce a new paradigm, programmable motion generation. In this paradigm, any given motion control task is broken down into a combination of atomic constraints. These constraints are then programmed into an error function that quantifies the degree to which a motion sequence adheres to them. We utilize a pretrained motion generation model and optimize its latent code to minimize the error function of the generated motion. Consequently, the generated motion not only inherits the prior of the generative model but also satisfies the requirements of the compounded constraints. Our experiments demonstrate that our approach can generate high-quality motions when addressing a wide range of unseen tasks. These tasks encompass motion control by motion dynamics, geometric constraints, physical laws, interactions with scenes, objects or the character's own body parts, etc. All of these are achieved in a unified approach, without the need for ad-hoc paired training data collection or specialized network designs. During the programming of novel tasks, we observed the emergence of new skills beyond those of the prior model. With the assistance of large language models, we also achieved automatic programming. We hope that this work will pave the way for the motion control of general AI agents. Xiaohang Zhan, Shaoli Huang, Tai-Jiang Mu, Ying Shan |
CVPR | 4 |
| 2024 | Theoretically Achieving Continuous Representation of Oriented Bounding BoxesabstractConsiderable efforts have been devoted to Oriented Ob-ject Detection (OOD). However, one lasting issue regarding the discontinuity in Oriented Bounding Box (OBB) rep-resentation remains unresolved, which is an inherent bot-tleneck for extant OOD methods. This paper endeavors to completely solve this issue in a theoretically guaranteed manner and puts an end to the ad-hoc efforts in this di-rection. Prior studies typically can only address one of the two cases of discontinuity: rotation and aspect ratio, and often inadvertently introduce decoding discontinuity, e.g. Decoding Incompleteness (DI) and Decoding Ambi-guity (DA) as discussed in literature. Specifically, we pro-pose a novel representation method called Continuous OBB (COBB), which can be readily integrated into existing de-tectors e.g. Faster-RCNN as a plugin. It can theoreti-cally ensure continuity in bounding box regression which to our best knowledge, has not been achieved in literature for rectangle-based object representation. For fairness and transparency of experiments, we have developed a modu-larized benchmark based on the open-source deep learning framework Jittor's detection toolbox JDetfor OOD evaluation. On the popular DOTA dataset, by integrating Faster-RCNN as the same baseline model, our new method out-performs the peer method Gliding Vertex by 1.13% mAP50(relative improvement 1.54%), and 2.46% mAP75(relative improvement 5.91%), without any tricks. Zi-Kai Xiao, Guo-Ye Yang, Xue Yang 0005, Tai-Jiang Mu, Junchi Yan, Shi-Min Hu 0001 |
CVPR | 4 |
| 2024 | Recovering Complete Actions for Cross-dataset Skeleton Action RecognitionabstractDespite huge progress in skeleton-based action recognition, its generalizability to different domains remains a challenging issue.
In this paper, to solve the skeleton action generalization problem, we present a recover-and-resample augmentation framework based on a novel complete action prior. We observe that human daily actions are confronted with temporal mismatch across different datasets, as they are usually partial observations of their complete action sequences. By recovering complete actions and resampling from these full sequences, we can generate strong augmentations for unseen domains. At the same time, we discover the nature of general action completeness within large datasets, indicated by the per-frame diversity over time. This allows us to exploit two assets of transferable knowledge that can be shared across action samples and be helpful for action completion: boundary poses for determining the action start, and linear temporal transforms for capturing global action patterns. Therefore, we formulate the recovering stage as a two-step stochastic action completion with boundary pose-conditioned extrapolation followed by smooth linear transforms. Both the boundary poses and linear transforms can be efficiently learned from the whole dataset via clustering. We validate our approach on a cross-dataset setting with three skeleton action datasets, outperforming other domain generalization approaches by a considerable margin. Yujiang Li, Tai-Jiang Mu, Shi-Min Hu 0001 |
NeurIPS | 3 |
| 2024 | Spoofing Transaction Detection with Group Perceptual Enhanced Graph Neural Network
Tai-Jiang Mu, Xiaodong Ning |
ECML/PKDD (9) | 2 |
| 2024 | EVSplitting: An Efficient and Visually Consistent Splitting Algorithm for 3D Gaussian SplattingabstractThis paper presents EVSplitting, an efficient and visually consistent splitting algorithm for 3D Gaussian Splatting (3DGS). It is designed to make operating 3DGS as easy and effective as other 3D explicit representations, readily for industrial productions. The challenges of above target are: 1) The huge number and complex attributes of 3DGS make it tough to explicitly operate on 3DGS in a real-time and learning-free manner; 2) The visual effect of 3DGS is very difficult to maintain during explicit operations and 3) The anisotropism of Gaussian always leads to blurs and artifacts. As far as we know, no prior work can address these challenges well. In this work, we introduce a direct and efficient 3DGS splitting algorithm to solve them. Specifically, we formulate the 3DGS splitting as two minimization problems that aim to ensure visual consistency and reduce Gaussian overflow across boundary (splitting plane), respectively. Firstly, we impose conservations on the zero-, first- and second-order moments of the weighted Gaussian distribution to guarantee visual consistency. Secondly, we reduce the boundary overflow with a special constraint on the aforementioned conservations. With these conservations and constraints, we derive a closed-form solution for the 3DGS splitting problem. This yields an easy-to-implement, plug-and-play, efficient and fundamental tool, benefiting various downstream applications of 3DGS. Qi-Yuan Feng, Geng-Chen Cao, Haoxiang Chen 0004, Qun-Ce Xu, Tai-Jiang Mu, Ralph R. Martin, Shi-Min Hu 0001 |
SIGGRAPH Asia | 5 |
| 2024 | DIScene: Object Decoupling and Interaction Modeling for Complex Scene Generation
Xiao-Lei Li, Haoxiang Chen 0004, Tai-Jiang Mu, Shi-Min Hu 0001 |
SIGGRAPH Asia | 4 |
| 2024 | FragmentDiff: A Diffusion Model for Fractured Object Assembly
Qun-Ce Xu, Haoxiang Chen 0004, Jiacheng Hua, Xiaohua Zhan, Yongliang Yang 0002, Tai-Jiang Mu |
SIGGRAPH Asia | 6 |
| 2024 | Sketch-2-4D: Sketch driven dynamic 3D scene generationabstractSketch-based content generation offers flexible controllability, making it a promising narrative avenue in film production. Directors often visualize their imagination by crafting storyboards using sketches and textual descriptions for each shot. However, current video generation methods suffer from three-dimensional inconsistencies, with notably artifacts during large motion or camera pans around scenes. A suitable solution is to directly generate 4D scene, enabling consistent dynamic three-dimensional scenes generation. We define the Sketch-2-4D problem, aiming to enhance controllability and consistency in this context. We propose a novel Control Score Distillation Sampling (SDS-C) for sketch-based 4D scene generation, providing precise control over scene dynamics. We further design Spatial Consistency Modules and Temporal Consistency Modules to tackle the temporal and spatial inconsistencies introduced by sketch-based control, respectively. Extensive experiments have demonstrated the effectiveness of our approach. Dong-Yu Chen, Tai-Jiang Mu |
Graph. Model. | 3 |
| 2024 | FilterGNN: Image feature matching with cascaded outlier filters and linear attentionabstractThe cross-view matching of local image features is a fundamental task in visual localization and 3D reconstruction. This study proposes FilterGNN, a transformer-based graph neural network (GNN), aiming to improve the matching efficiency and accuracy of visual descriptors. Based on high matching sparseness and coarse-to-fine covisible area detection, FilterGNN utilizes cascaded optimal graph-matching filter modules to dynamically reject outlier matches. Moreover, we successfully adapted linear attention in FilterGNN with post-instance normalization support, which significantly reduces the complexity of complete graph learning from O ( N 2 ) to O ( N ). Experiments show that FilterGNN requires only 6% of the time cost and 33.3% of the memory cost compared with SuperGlue under a large-scale input size and achieves a competitive performance in various tasks, such as pose estimation, visual localization, and sparse 3D reconstruction. Junxiong Cai, Tai-Jiang Mu, Yukun Lai |
Comput. Vis. Media | 2 |
| 2024 | Multi3D: 3D-aware multimodal image synthesisabstract3D-aware image synthesis has attained high quality and robust 3D consistency. Existing 3D controllable generative models are designed to synthesize 3D-aware images through a single modality, such as 2D segmentation or sketches, but lack the ability to finely control generated content, such as texture and age. In pursuit of enhancing user-guided controllability, we propose Multi3D, a 3D-aware controllable image synthesis model that supports multi-modal input. Our model can govern the geometry of the generated image using a 2D label map, such as a segmentation or sketch map, while concurrently regulating the appearance of the generated image through a textual description. To demonstrate the effectiveness of our method, we have conducted experiments on multiple datasets, including CelebAMask-HQ, AFHQ-cat, and shapenet-car. Qualitative and quantitative evaluations show that our method outperforms existing state-of-the-art methods. Wenyang Zhou, Tai-Jiang Mu |
Comput. Vis. Media | 3 |
| 2024 | Tuning Vision-Language Models With Multiple Prototypes ClusteringabstractBenefiting from advances in large-scale pre-training, foundation models, have demonstrated remarkable capability in the fields of natural language processing, computer vision, among others. However, to achieve expert-level performance in specific applications, such models often need to be fine-tuned with domain-specific knowledge. In this paper, we focus on enabling vision-language models to unleash more potential for visual understanding tasks under few-shot tuning. Specifically, we propose a novel adapter, dubbed as lusterAdapter, which is based on trainable multiple prototypes clustering algorithm, for tuning the CLIP model. It can not only alleviate the concern of catastrophic forgetting of foundation models by introducing anchors to inherit common knowledge, but also improve the utilization efficiency of few annotated samples via bringing in clustering and domain priors, thereby improving the performance of few-shot tuning. We have conducted extensive experiments on 11 common classification benchmarks. The results show our method significantly surpasses the original CLIP and achieves state-of-the-art (SOTA) performance under all benchmarks and settings. For example, under the 16-shot setting, our method exhibits a remarkable improvement over the original CLIP by 19.6%, and also surpasses TIP-Adapter and GraphAdapter by 2.7% and 2.2%, respectively, in terms of average accuracy across the 11 benchmarks. Menghao Guo 0001, Yi Zhang 0099, Tai-Jiang Mu, Sharon X. Huang, Shi-Min Hu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Learning Virtual View Selection for 3D Scene Semantic Segmentationabstract2D-3D joint learning is essential and effective for fundamental 3D vision tasks, such as 3D semantic segmentation, due to the complementary information these two visual modalities contain. Most current 3D scene semantic segmentation methods process 2D images "as they are", i.e., only real captured 2D images are used. However, such captured 2D images may be redundant, with abundant occlusion and/or limited field of view (FoV), leading to poor performance for the current methods involving 2D inputs. In this paper, we propose a general learning framework for joint 2D-3D scene understanding by selecting informative virtual 2D views of the underlying 3D scene. We then feed both the 3D geometry and the generated virtual 2D views into any joint 2D-3D-input or pure 3D-input based deep neural models for improving 3D scene understanding. Specifically, we generate virtual 2D views based on an information score map learned from the current 3D scene semantic segmentation results. To achieve this, we formalize the learning of the information score map as a deep reinforcement learning process, which rewards good predictions using a deep neural network. To obtain a compact set of virtual 2D views that jointly cover informative surfaces of the 3D scene as much as possible, we further propose an efficient greedy virtual view coverage strategy in the normal-sensitive 6D space, including 3-dimensional point coordinates and 3-dimensional normal. We have validated our proposed framework for various joint 2D-3D-input or pure 3D-input based deep neural models on two real-world 3D scene datasets, i.e., ScanNet v2 and S3DIS, and the results demonstrate that our method obtains a consistent gain over baseline models and achieves new top accuracy for joint 2D and 3D scene semantic segmentation. Code is available at https://github.com/smy-THU/VirtualViewSelection. Tai-Jiang Mu, Ming-Yuan Shen, Yukun Lai, Shi-Min Hu 0001 |
IEEE Trans. Image Process. | 1 |
| 2024 | Mesh Neural Networks Based on Dual Graph PyramidsabstractDeep neural networks (DNNs) have been widely used for mesh processing in recent years. However, current DNNs can not process arbitrary meshes efficiently. On the one hand, most DNNs expect 2-manifold, watertight meshes, but many meshes, whether manually designed or automatically generated, may have gaps, non-manifold geometry, or other defects. On the other hand, the irregular structure of meshes also brings challenges to building hierarchical structures and aggregating local geometric information, which is critical to conduct DNNs. In this paper, we present DGNet, an efficient, effective and generic deep neural mesh processing network based on dual graph pyramids; it can handle arbitrary meshes. First, we construct dual graph pyramids for meshes to guide feature propagation between hierarchical levels for both downsampling and upsampling. Second, we propose a novel convolution to aggregate local features on the proposed hierarchical graphs. By utilizing both geodesic neighbors and euclidean neighbors, the network enables feature aggregation both within local surface patches and between isolated mesh components. Experimental results demonstrate that DGNet can be applied to both shape analysis and large-scale scene understanding. Furthermore, it achieves superior performance on various benchmarks, including ShapeNetCore, HumanBody, ScanNet and Matterport3D. Code and models will be available at https://github.com/li-xl/DGNet. Xiang-Li Li, Zheng-Ning Liu, Tuo Chen, Tai-Jiang Mu, Ralph R. Martin, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | Generating animatable 3D cartoon faces from single portraitsabstractBackground \nWith the development of virtual reality (VR) technology, there is a growing need for customized 3D avatars. However, traditional methods for 3D avatar modeling are either time-consuming or fail to retain the similarity to the person being modeled. This study presents a novel framework for generating animatable 3D cartoon faces from a single portrait image. \n \nMethods \nFirst, we transferred an input real-world portrait to a stylized cartoon image using StyleGAN. We then proposed a two-stage reconstruction method to recover a 3D cartoon face with detailed texture. Our two-stage strategy initially performs coarse estimation based on template models and subsequently refines the model by nonrigid deformation under landmark supervision. Finally, we proposed a semantic-preserving face-rigging method based on manually created templates and deformation transfer. \n \nConclusions \nCompared with prior arts, the qualitative and quantitative results show that our method achieves better accuracy, aesthetics, and similarity criteria. Furthermore, we demonstrated the capability of the proposed 3D model for real-time facial animation. Chuanyu Pan, Tai-Jiang Mu, Yukun Lai |
Virtual Real. Intell. Hardw. | 3 |
| 2023 | Conspiracy Spoofing Orders Detection with Transformer-Based Deep Graph Learning
Tai-Jiang Mu, Xiaodong Ning |
ADMA (2) | 2 |
| 2023 | Long Range Pooling for 3D Large-Scale Scene UnderstandingabstractInspired by the success of recent vision transformers and large kernel design in convolutional neural networks (CNNs), in this paper, we analyze and explore essential reasons for their success. We claim two factors that are critical for 3D large-scale scene understanding: a larger receptive field and operations with greater non-linearity. The former is responsible for providing long range contexts and the latter can enhance the capacity of the network. To achieve the above properties, we propose a simple yet effective long range pooling (LRP) module using dilation max pooling, which provides a network with a large adaptive receptive field. LRP has few parameters, and can be readily added to current CNNs. Also, based on LRP, we present an entire network architecture, LRPNet, for 3D understanding. Ablation studies are presented to support our claims, and show that the LRP module achieves better results than large kernel convolution yet with reduced computation, due to its non-linearity. We also demonstrate the superiority of LRPNet on various benchmarks: LRPNet performs the best on ScanNet and surpasses other CNN-based methods on S3DIS and Matterport3D. Code will be avalible at https://github.com/li-xl/LRPNet. Xiang-Li Li, Menghao Guo 0001, Tai-Jiang Mu, Ralph R. Martin, Shi-Min Hu 0001 |
CVPR | 3 |
| 2023 | Hierarchical Transformer-based Siamese Network for Related Trading Detection in Financial MarketabstractThe phenomenon of related trading, in which organized communities engage in coordinated trading activities, poses a significant threat to the financial markets. Such activities can facilitate financial fraud, such as insider trading, price control, and market manipulation. Therefore, the detection of related trading is crucial for regulators to take actions to maintain market fairness and reduce market risk. The key challenge in detecting related trading is to measure the relevance of the trading behaviors of traders. Trade data is often represented as time series data, and traditional methods for analyzing such data typically focus on designing complex handcrafted features to measure the similarity of these time series. However, these methods are often incapable of capturing the complexity and variability of trading behaviors, which can be confusing and misleading. In this paper, we address this limitation by introducing a Hierarchical Transformer-based Siamese Network (HTSN) for related trading detection. The HTSN learns the correlation of the trade data from two accounts in an end-to-end manner, and is able to better capture trading information by splitting the data into different scales and stacking multiple Transformer encoders to extract features hierarchically. The experimental results on real-world data from China's futures market, indicate that the proposed HTSN model substantially outperforms previous approaches in detecting related trading. Tai-Jiang Mu, Guoping Zhao |
IJCNN | 2 |
| 2023 | MWFormer: Mesh Understanding with Window-based Transformer
Hao-Yang Peng, Menghao Guo 0001, Zheng-Ning Liu, Yongliang Yang 0002, Tai-Jiang Mu |
Comput. Graph. | 5 |
| 2023 | Neural 3D reconstruction from sparse views using geometric priorsabstractSparse view 3D reconstruction has attracted increasing attention with the development of neural implicit 3D representation. Existing methods usually only make use of 2D views, requiring a dense set of input views for accurate 3D reconstruction. In this paper, we show that accurate 3D reconstruction can be achieved by incorporating geometric priors into neural implicit 3D reconstruction. Our method adopts the signed distance function as the 3D representation, and learns a generalizable 3D surface reconstruction model from sparse views. Specifically, we build a more effective and sparse feature volume from the input views by using corresponding depth maps, which can be provided by depth sensors or directly predicted from the input views. We recover better geometric details by imposing both depth and surface normal constraints in addition to the color loss when training the neural implicit 3D representation. Experiments demonstrate that our method both outperforms state-of-the-art approaches, and achieves good generalizability. Tai-Jiang Mu, Haoxiang Chen 0004, Junxiong Cai |
Comput. Vis. Media | 1 |
| 2023 | A survey of deep learning-based 3D shape generationabstractDeep learning has been successfully used for tasks in the 2D image domain. Research on 3D computer vision and deep geometry learning has also attracted attention. Considerable achievements have been made regarding feature extraction and discrimination of 3D shapes. Following recent advances in deep generative models such as generative adversarial networks, effective generation of 3D shapes has become an active research topic. Unlike 2D images with a regular grid structure, 3D shapes have various representations, such as voxels, point clouds, meshes, and implicit functions. For deep learning of 3D shapes, shape representation has to be taken into account as there is no unified representation that can cover all tasks well. Factors such as the representativeness of geometry and topology often largely affect the quality of the generated 3D shapes. In this survey, we comprehensively review works on deep-learning-based 3D shape generation by classifying and discussing them in terms of the underlying shape representation and the architecture of the shape generator. The advantages and disadvantages of each class are further analyzed. We also consider the 3D shape datasets commonly used for shape generation. Finally, we present several potential research directions that hopefully can inspire future works on this topic. Qun-Ce Xu, Tai-Jiang Mu, Yongliang Yang 0002 |
Comput. Vis. Media | 2 |
| 2023 | Beyond Self-Attention: External Attention Using Two Linear Layers for Visual TasksabstractAttention mechanisms, especially self-attention, have played an increasingly important role in deep feature representation for visual tasks. Self-attention updates the feature at each position by computing a weighted sum of features using pair-wise affinities across all positions to capture the long-range dependency within a single sample. However, self-attention has quadratic complexity and ignores potential correlation between different samples. This article proposes a novel attention mechanism which we call external attention, based on two external, small, learnable, shared memories, which can be implemented easily by simply using two cascaded linear layers and two normalization layers; it conveniently replaces self-attention in existing popular architectures. External attention has linear complexity and implicitly considers the correlations between all data samples. We further incorporate the multi-head mechanism into external attention to provide an all-MLP architecture, external attention MLP (EAMLP), for image classification. Extensive experiments on image classification, object detection, semantic segmentation, instance segmentation, image generation, and point cloud analysis reveal that our method provides results comparable or superior to the self-attention mechanism and some of its variants, with much lower computational and memory costs. Menghao Guo 0001, Zheng-Ning Liu, Tai-Jiang Mu, Shi-Min Hu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | SPS: Accurate and Real-Time Semantic Positioning System Based on Low-Cost DEM MapsabstractThis paper presents a Semantic Positioning System (SPS) to enhance the accuracy of mobile device geo-localization in outdoor urban environments. Although the traditional Global Positioning System (GPS) can offer a rough localization, it lacks the necessary accuracy for applications such as Augmented Reality (AR). Our SPS integrates Geographic Information System (GIS) data, GPS signals, and visual image information to estimate the 6 Degree-of-Freedom (DoF) pose through cross-view semantic matching. This approach has excellent scalability to support GIS context with Levels of Detail (LOD). The map data representation is Digital Elevation Model (DEM), a cost-effective aerial map that allows for fast deployment for large-scale areas. However, the DEM lacks geometric and texture details, making it challenging for traditional visual feature extraction to establish pixel/voxel level cross-view correspondences. To address this, we sample observation pixels from the query ground-view image using predicted semantic labels. We then propose an iterative homography estimation method with semantic correspondences. To improve the efficiency of the overall system, we further employ a heuristic search to speedup the matching process. The proposed method is robust, real-time, and automatic. Quantitative experiments on the challenging Bund dataset show that we achieve a positioning accuracy of 73.24%, surpassing the baseline skyline-based method by 20%. Compared with the state-of-the-art semantic-based approach on the Kitti dataset, we improve the positioning accuracy by an average of 5%. Junxiong Cai, Wensen Feng, Haoxiang Chen 0004, Tai-Jiang Mu |
IEEE Trans. Image Process. | 4 |
| 2023 | Skeleton-CutMix: Mixing Up Skeleton With Probabilistic Bone Exchange for Supervised Domain AdaptationabstractWe present Skeleton-CutMix, a simple and effective skeleton augmentation framework for supervised domain adaptation and show its advantage in skeleton-based action recognition tasks. Existing approaches usually perform domain adaptation for action recognition with elaborate loss functions that aim to achieve domain alignment. However, they fail to capture the intrinsic characteristics of skeleton representation. Benefiting from the well-defined correspondence between bones of a pair of skeletons, we instead mitigate domain shift by fabricating skeleton data in a mixed domain, which mixes up bones from the source domain and the target domain. The fabricated skeletons in the mixed domain can be used to augment training data and train a more general and robust model for action recognition. Specifically, we hallucinate new skeletons by using pairs of skeletons from the source and target domains; a new skeleton is generated by exchanging some bones from the skeleton in the source domain with corresponding bones from the skeleton in the target domain, which resembles a cut-and-mix operation. When exchanging bones from different domains, we introduce a class-specific bone sampling strategy so that bones that are more important for an action class are exchanged with higher probability when generating augmentation samples for that class. We show experimentally that the simple bone exchange strategy for augmentation is efficient and effective and that distinctive motion features are preserved while mixing both action and style across domains. We validate our method in cross-dataset and cross-age settings on NTU-60 and ETRI-Activity3D datasets with an average gain of over 3% in terms of action recognition accuracy, and demonstrate its superior performance over previous domain adaptation approaches as well as other skeleton augmentation strategies. Yuhe Liu, Tai-Jiang Mu, Sharon X. Huang, Shi-Min Hu 0001 |
IEEE Trans. Image Process. | 3 |
| 2023 | Sampling Equivariant Self-Attention Networks for Object Detection in Aerial ImagesabstractObjects in aerial images show greater variations in scale and orientation than in other images, making them harder to detect using vanilla deep convolutional neural networks. Networks with sampling equivariance can adapt sampling from input feature maps to object transformation, allowing a convolutional kernel to extract effective object features under different transformations. However, methods such as deformable convolutional networks can only provide sampling equivariance under certain circumstances, as they sample by location. We propose sampling equivariant self-attention networks, which treat self-attention restricted to a local image patch as convolution sampling by masks instead of locations, and a transformation embedding module to improve the equivariant sampling further. We further propose a novel randomized normalization module to enhance network generalization and a quantitative evaluation metric to fairly evaluate the ability of sampling equivariance of different models. Experiments show that our model provides significantly better sampling equivariance than existing methods without additional supervision and can thus extract more effective image features. Our model achieves state-of-the-art results on the DOTA-v1.0, DOTA-v1.5, and HRSC2016 datasets without additional computations or parameters. Guo-Ye Yang, Xiang-Li Li, Zi-Kai Xiao, Tai-Jiang Mu, Ralph R. Martin, Shi-Min Hu 0001 |
IEEE Trans. Image Process. | 4 |
| 2023 | Recursive-NeRF: An Efficient and Dynamically Growing NeRFabstractView synthesis methods using implicit continuous shape representations learned from a set of images, such as the Neural Radiance Field (NeRF) method, have gained increasing attention due to their high quality imagery and scalability to high resolution. However, the heavy computation required by its volumetric approach prevents NeRF from being useful in practice; minutes are taken to render a single image of a few megapixels. Now, an image of a scene can be rendered in a level-of-detail manner, so we posit that a complicated region of the scene should be represented by a large neural network while a small neural network is capable of encoding a simple region, enabling a balance between efficiency and quality. Recursive-NeRF is our embodiment of this idea, providing an efficient and adaptive rendering and training approach for NeRF. The core of Recursive-NeRF learns uncertainties for query coordinates, representing the quality of the predicted color and volumetric intensity at each level. Only query coordinates with high uncertainties are forwarded to the next level to a bigger neural network with a more powerful representational capability. The final rendered image is a composition of results from neural networks of all levels. Our evaluation on public datasets and a large-scale scene dataset we collected shows that Recursive-NeRF is more efficient than NeRF while providing state-of-the-art quality. The code will be available at https://github.com/Gword/Recursive-NeRF. Wenyang Zhou, Hao-Yang Peng, Dun Liang, Tai-Jiang Mu, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2023 | SceneViewer: Automating Residential Photography in Virtual EnvironmentsabstractSelecting views is one of the most common but overlooked procedures in topics related to 3D scenes. Typically, existing applications and researchers manually select views through a trial-and-error process or "preset" a direction, such as the top-down views. For example, literature for scene synthesis requires views for visualizing scenes. Research on panorama and VR also require initial placements for cameras, etc. This article presents SceneViewer, an integrated system for automatic view selections. Our system is achieved by applying rules of interior photography, which guides potential views and seeks better views. Through experiments and applications, we show the potentiality and novelty of the proposed method. Shao-Kui Zhang, Hou Tam, Yi-Xiao Li, Tai-Jiang Mu, Song-Hai Zhang |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2022 | ClusterGNN: Cluster-based Coarse-to-Fine Graph Neural Network for Efficient Feature MatchingabstractGraph Neural Networks (GNNs) with attention have been successfully applied for learning visual feature matching. However, current methods learn with complete graphs, resulting in a quadratic complexity in the number of features. Motivated by a prior observation that self- and cross- attention matrices converge to a sparse representation, we propose ClusterGNN, an attentional GNN architecture which operates on clusters for learning the feature matching task. Using a progressive clustering module we adaptively divide keypoints into different subgraphs to reduce redundant connectivity, and employ a coarse-to-fine paradigm for mitigating miss-classification within images. Our approach yields a 59.7% reduction in runtime and 58.4% reduction in memory consumption for dense detection, compared to current state-of-the-art GNN-based matching, while achieving a competitive performance on various computer vision tasks. Junxiong Cai, Yoli Shavit, Tai-Jiang Mu, Wensen Feng, Kai Zhang 0012 |
CVPR | 4 |
| 2022 | CIRCLE: Convolutional Implicit Reconstruction and Completion for Large-Scale Indoor Scene
Haoxiang Chen 0004, Tai-Jiang Mu, Shi-Min Hu 0001 |
ECCV (32) | 3 |
| 2022 | Joint Hand and Object Pose Estimation from a Single RGB Image using High-level 2D ConstraintsabstractAbstract Joint pose estimation of human hands and objects from a single RGB image is an important topic for AR/VR, robot manipulation, etc. It is common practice to determine both poses directly from the image; some recent methods attempt to improve the initial poses using a variety of contact‐based approaches. However, few methods take the real physical constraints conveyed by the image into consideration, leading to less realistic results than the initial estimates. To overcome this problem, we make use of a set of high‐level 2D features which can be directly extracted from the image in a new pipeline which combines contact approaches and these constraints during optimization. Our pipeline achieves better results than direct regression or contact‐based optimization: they are closer to the ground truth and provide high quality contact. Haoxuan Song, Tai-Jiang Mu, Ralph R. Martin |
Comput. Graph. Forum | 2 |
| 2022 | ObjectFusion: Accurate object-level SLAM with neural object priors
Zixin Zou, Shi-Sheng Huang, Tai-Jiang Mu, Yu-Ping Wang 0001 |
Graph. Model. | 3 |
| 2022 | Attention mechanisms in computer vision: A surveyabstractHumans can naturally and effectively find salient regions in complex scenes. Motivated by this observation, attention mechanisms were introduced into computer vision with the aim of imitating this aspect of the human visual system. Such an attention mechanism can be regarded as a dynamic weight adjustment process based on features of the input image. Attention mechanisms have achieved great success in many visual tasks, including image classification, object detection, semantic segmentation, video understanding, image generation, 3D vision, multimodal tasks, and self-supervised learning. In this survey, we provide a comprehensive review of various attention mechanisms in computer vision and categorize them according to approach, such as channel attention, spatial attention, temporal attention, and branch attention; a related repository https://github.com/MenghaoGuo/Awesome-Vision-Attentions is dedicated to collecting related work. We also suggest future directions for attention mechanism research. Menghao Guo 0001, Tian-Xing Xu, Jiang-Jiang Liu 0001, Zheng-Ning Liu, Peng-Tao Jiang, Tai-Jiang Mu, Song-Hai Zhang, Ralph R. Martin, Ming-Ming Cheng, Shi-Min Hu 0001 |
Comput. Vis. Media | 6 |
| 2022 | Subdivision-based Mesh Convolution NetworksabstractConvolutionalneural networks (CNNs) have made great breakthroughs in two-dimensional (2D) computer vision. However, their irregular structure makes it hard to harness the potential of CNNs directly on meshes. A subdivision surface provides a hierarchical multi-resolution structure in which each face in a closed 2-manifold triangle mesh is exactly adjacent to three faces. Motivated by these two observations, this article presents SubdivNet , an innovative and versatile CNN framework for three-dimensional (3D) triangle meshes with Loop subdivision sequence connectivity. Making an analogy between mesh faces and pixels in a 2D image allows us to present a mesh convolution operator to aggregate local features from nearby faces. By exploiting face neighborhoods, this convolution can support standard 2D convolutional network concepts, e.g., variable kernel size, stride, and dilation. Based on the multi-resolution hierarchy, we make use of pooling layers that uniformly merge four faces into one and an upsampling method that splits one face into four. Thereby, many popular 2D CNN architectures can be easily adapted to process 3D meshes. Meshes with arbitrary connectivity can be remeshed to have Loop subdivision sequence connectivity via self-parameterization, making SubdivNet a general approach. Extensive evaluation and various applications demonstrate SubdivNet’s effectiveness and efficiency. Shi-Min Hu 0001, Zheng-Ning Liu, Menghao Guo 0001, Junxiong Cai, Tai-Jiang Mu, Ralph R. Martin |
ACM Trans. Graph. | 6 |
| 2022 | Accurate Dynamic SLAM Using CRF-Based Long-Term ConsistencyabstractAccurate camera pose estimation is essential and challenging for real world dynamic 3D reconstruction and augmented reality applications. In this article, we present a novel RGB-D SLAM approach for accurate camera pose tracking in dynamic environments. Previous methods detect dynamic components only across a short time-span of consecutive frames. Instead, we provide a more accurate dynamic 3D landmark detection method, followed by the use of long-term consistency via conditional random fields, which leverages long-term observations from multiple frames. Specifically, we first introduce an efficient initial camera pose estimation method based on distinguishing dynamic from static points using graph-cut RANSAC. These static/dynamic labels are used as priors for the unary potential in the conditional random fields, which further improves the accuracy of dynamic 3D landmark detection. Evaluation using the TUM and Bonn RGB-D dynamic datasets shows that our approach significantly outperforms state-of-the-art methods, providing much more accurate camera trajectory estimation in a variety of highly dynamic environments. We also show that dynamic 3D reconstruction can benefit from the camera poses estimated by our RGB-D SLAM approach. Zheng-Jun Du, Shi-Sheng Huang, Tai-Jiang Mu, Qunhe Zhao, Ralph R. Martin, Kun Xu 0003 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2021 | LinkNet: 2D-3D linked multi-modal network for online semantic segmentation of RGB-D videos
Junxiong Cai, Tai-Jiang Mu, Yukun Lai, Shi-Min Hu 0001 |
Comput. Graph. | 2 |
| 2021 | PCT: Point cloud transformerabstractThe irregular domain and lack of ordering make it challenging to design deep neural networks for point cloud processing. This paper presents a novel framework named Point Cloud Transformer (PCT) for point cloud learning. PCT is based on Transformer, which achieves huge success in natural language processing and displays great potential in image processing. It is inherently permutation invariant for processing a sequence of points, making it well-suited for point cloud learning. To better capture local context within the point cloud, we enhance input embedding with the support of farthest point sampling and nearest neighbor search. Extensive experiments demonstrate that the PCT achieves the state-of-the-art performance on shape classification, part segmentation, semantic segmentation, and normal estimation tasks. Menghao Guo 0001, Junxiong Cai, Zheng-Ning Liu, Tai-Jiang Mu, Ralph R. Martin, Shi-Min Hu 0001 |
Comput. Vis. Media | 4 |
| 2021 | Can attention enable MLPs to catch up with CNNs?abstractIn the first week of May 2021, researchers from four different institutions: Google, Menghao Guo 0001, Zheng-Ning Liu, Tai-Jiang Mu, Dun Liang, Ralph R. Martin, Shi-Min Hu 0001 |
Comput. Vis. Media | 3 |
| 2021 | Detecting human - object interaction with multi-level pairwise feature networkabstractHuman–object interaction (HOI) detection is crucial for human-centric image understanding which aims to infer ⟨human, action, object⟩ triplets within an image. Recent studies often exploit visual features and the spatial configuration of a human–object pair in order to learn the action linking the human and object in the pair. We argue that such a paradigm of pairwise feature extraction and action inference can be applied not only at the whole human and object instance level, but also at the part level at which a body part interacts with an object, and at the semantic level by considering the semantic label of an object along with human appearance and human–object spatial configuration, to infer the action. We thus propose a multi-level pairwise feature network (PFNet) for detecting human–object interactions. The network consists of three parallel streams to characterize HOI utilizing pairwise features at the above three levels; the three streams are finally fused to give the action prediction. Extensive experiments show that our proposed PFNet outperforms other state-of-the-art methods on the V-COCO dataset and achieves comparable results to the state-of-the-art on the HICO-DET dataset. Tai-Jiang Mu, Sharon X. Huang |
Comput. Vis. Media | 2 |
| 2021 | HDR-Net-Fusion: Real-time 3D dynamic scene reconstruction with a hierarchical deep reinforcement networkabstractAbstract Reconstructing dynamic scenes with commodity depth cameras has many applications in computer graphics, computer vision, and robotics. However, due to the presence of noise and erroneous observations from data capturing devices and the inherently ill-posed nature of non-rigid registration with insufficient information, traditional approaches often produce low-quality geometry with holes, bumps, and misalignments. We propose a novel 3D dynamic reconstruction system, named HDR-Net-Fusion, which learns to simultaneously reconstruct and refine the geometry on the fly with a sparse embedded deformation graph of surfels, using a hierarchical deep reinforcement (HDR) network. The latter comprises two parts: a global HDR-Net which rapidly detects local regions with large geometric errors, and a local HDR-Net serving as a local patch refinement operator to promptly complete and enhance such regions. Training the global HDR-Net is formulated as a novel reinforcement learning problem to implicitly learn the region selection strategy with the goal of improving the overall reconstruction quality. The applicability and efficiency of our approach are demonstrated using a large-scale dynamic reconstruction dataset. Our method can reconstruct geometry with higher quality than traditional methods. Haoxuan Song, Yan-Pei Cao 0001, Tai-Jiang Mu |
Comput. Vis. Media | 4 |
| 2021 | Supervoxel Convolution for Online 3D Semantic SegmentationabstractOnline 3D semantic segmentation, which aims to perform real-time 3D scene reconstruction along with semantic segmentation, is an important but challenging topic. A key challenge is to strike a balance between efficiency and segmentation accuracy. There are very few deep-learning-based solutions to this problem, since the commonly used deep representations based on volumetric-grids or points do not provide efficient 3D representation and organization structure for online segmentation. Observing that on-surface supervoxels, i.e., clusters of on-surface voxels, provide a compact representation of 3D surfaces and brings efficient connectivity structure via supervoxel clustering, we explore a supervoxel-based deep learning solution for this task. To this end, we contribute a novel convolution operation (SVConv) directly on supervoxels. SVConv can efficiently fuse the multi-view 2D features and 3D features projected on supervoxels during the online 3D reconstruction, and leads to an effective supervoxel-based convolutional neural network, termed as Supervoxel-CNN , enabling 2D-3D joint learning for 3D semantic prediction. With the Supervoxel-CNN , we propose a clustering-then-prediction online 3D semantic segmentation approach. The extensive evaluations on the public 3D indoor scene datasets show that our approach significantly outperforms the existing online semantic segmentation systems in terms of efficiency or accuracy. Shi-Sheng Huang, Tai-Jiang Mu, Hongbo Fu 0001, Shi-Min Hu 0001 |
ACM Trans. Graph. | 3 |
| 2020 | ClusterVO: Clustering Moving Instances and Estimating Visual Odometry for Self and SurroundingsabstractWe present ClusterVO, a stereo Visual Odometry which simultaneously clusters and estimates the motion of both ego and surrounding rigid clusters/objects. Unlike previous solutions relying on batch input or imposing priors on scene structure or dynamic object models, ClusterVO is online, general and thus can be used in various scenarios including indoor scene understanding and autonomous driving. At the core of our system lies a multi-level probabilistic association mechanism and a heterogeneous Conditional Random Field (CRF) clustering approach combining semantic, spatial and motion information to jointly infer cluster segmentations online for every frame. The poses of camera and dynamic objects are instantly solved through a sliding-window optimization. Our system is evaluated on Oxford Multimotion and KITTI dataset both quantitatively and qualitatively, reaching comparable results to state-of-the-art solutions on both odometry and dynamic trajectory recovery. Sheng Yang 0007, Tai-Jiang Mu, Shi-Min Hu 0001 |
CVPR | 3 |
| 2020 | Lidar-Monocular Visual Odometry using Point and Line FeaturesabstractWe introduce a novel lidar-monocular visual odometry approach using point and line features. Compared to previous point-only based lidar-visual odometry, our approach leverages more environment structure information by introducing both point and line features into pose estimation. We provide a robust method for point and line depth extraction, and formulate the extracted depth as prior factors for point-line bundle adjustment. This method greatly reduces the features' 3D ambiguity and thus improves the pose estimation accuracy. Besides, we also provide a purely visual motion tracking method and a novel scale correction scheme, leading to an efficient lidar-monocular visual odometry system with high accuracy. The evaluations on the public KITTI odometry benchmark show that our technique achieves more accurate pose estimation than the state-of-the-art approaches, and is sometimes even better than those leveraging semantic information. Shi-Sheng Huang, Tai-Jiang Mu, Hongbo Fu 0001, Shi-Min Hu 0001 |
ICRA | 3 |
| 2020 | WallNet: Reconstructing General Room Layouts from RGB Images
Zheng-Fei Kuang, Tai-Jiang Mu |
Graph. Model. | 4 |
| 2020 | S4Net: Single stage salient-instance segmentationabstractIn this paper, we consider salient instance segmentation. As well as producing bounding boxes, our network also outputs high-quality instance-level segments as initial selections to indicate the regions of interest. Taking into account the category-independent property of each target, we design a single stage salient instance segmentation framework, with a novel segmentation branch. Our new branch regards not only local context inside each detection window but also the surrounding context, enabling us to distinguish instances in the same scope even with partial occlusion. Our network is end-to-end trainable and is fast (running at 40 fps for images with resolution 320 × 320). We evaluate our approach on a publicly available benchmark and show that it outperforms alternative solutions. We also provide a thorough analysis of our design choices to help readers better understand the function of each part of our network. Source code can be found at https://github.com/RuochenFan/S4Net. Ruochen Fan, Ming-Ming Cheng, Qibin Hou, Tai-Jiang Mu, Jingdong Wang 0001, Shi-Min Hu 0001 |
Comput. Vis. Media | 4 |
| 2020 | A new dataset of dog breed images and a benchmark for finegrained classificationabstractIn this paper, we introduce an image dataset for fine-grained classification of dog breeds: the Tsinghua Dogs Dataset. It is currently the largest dataset for fine-grained classification of dogs, including 130 dog breeds and 70,428 real-world images. It has only one dog in each image and provides annotated bounding boxes for the whole body and head. In comparison to previous similar datasets, it contains more breeds and more carefully chosen images for each breed. The diversity within each breed is greater, with between 200 and 7000+ images for each breed. Annotation of the whole body and head makes the dataset not only suitable for the improvement of finegrained image classification models based on overall features, but also for those locating local informative parts. We show that dataset provides a tough challenge by benchmarking several state-of-the-art deep neural models. The dataset is available for academic purposes at https://cg.cs.tsinghua.edu.cn/ThuDogs/ . Ding-Nan Zou, Song-Hai Zhang, Tai-Jiang Mu |
Comput. Vis. Media | 3 |
| 2020 | Lane Detection: A Survey with New Results
Dun Liang, Shao-Kui Zhang, Tai-Jiang Mu, Sharon X. Huang |
J. Comput. Sci. Technol. | 4 |
| 2019 | S4Net: Single Stage Salient-Instance SegmentationabstractWe consider an interesting problem---salient instance segmentation. Other than producing approximate bounding boxes, our network also outputs high-quality instance-level segments. Taking into account the category-independent property of each target, we design a single stage salient instance segmentation framework, with a novel segmentation branch. Our new branch regards not only local context inside each detection window but also its surrounding context, enabling us to distinguish the instances in the same scope even with obstruction. Our network is end-to-end trainable and runs at a fast speed (40 fps when processing an image with resolution 320 x 320). We evaluate our approach on a public available benchmark and show that it outperforms other alternative solutions. We also provide a thorough analysis of the design choices to help readers better understand the functions of each part of our network. The source code can be found at https://github.com/RuochenFan/S4Net. Ruochen Fan, Ming-Ming Cheng, Qibin Hou, Tai-Jiang Mu, Jingdong Wang 0001, Shi-Min Hu 0001 |
CVPR | 4 |
| 2019 | Deep point-based scene labeling with depth mapping and geometric patch feature encoding
Junxiong Cai, Tai-Jiang Mu, Yukun Lai, Shi-Min Hu 0001 |
Graph. Model. | 2 |
| 2019 | SpinNet: Spinning convolutional network for lane boundary detectionabstractIn this paper, we propose a simple but effective framework for lane boundary detection, called SpinNet. Considering that cars or pedestrians often occlude lane boundaries and that the local features of lane boundaries are not distinctive, therefore, analyzing and collecting global context information is crucial for lane boundary detection. To this end, we design a novel spinning convolution layer and a brand-new lane parameterization branch in our network to detect lane boundaries from a global perspective. To extract features in narrow strip-shaped fields, we adopt strip-shaped convolutions with kernels which have 1 × n or n × 1 shape in the spinning convolution layer. To tackle the problem of that straight strip-shaped convolutions are only able to extract features in vertical or horizontal directions, we introduce the concept of feature map rotation to allow the convolutions to be applied in multiple directions so that more information can be collected concerning a whole lane boundary. Moreover, unlike most existing lane boundary detectors, which extract lane boundaries from segmentation masks, our lane boundary parameterization branch predicts a curve expression for the lane boundary for each pixel in the output feature map. And the network utilizes this information to predict the weights of the curve, to better form the final lane boundaries. Our framework is easy to implement and end-to-end trainable. Experiments show that our proposed SpinNet outperforms state-of-the-art methods. Ruochen Fan, Xuanrun Wang, Qibin Hou, Tai-Jiang Mu |
Comput. Vis. Media | 5 |
| 2019 | A Large Chinese Text Dataset in the Wild
Tailing Yuan, Zhe Zhu, Kun Xu 0003, Cheng-Jun Li, Tai-Jiang Mu, Shi-Min Hu 0001 |
J. Comput. Sci. Technol. | 5 |
| 2018 | Deep Video Stabilization Using Adversarial NetworksabstractAbstract Video stabilization is necessary for many hand‐held shot videos. In the past decades, although various video stabilization methods were proposed based on the smoothing of 2D, 2.5D or 3D camera paths, hardly have there been any deep learning methods to solve this problem. Instead of explicitly estimating and smoothing the camera path, we present a novel online deep learning framework to learn the stabilization transformation for each unsteady frame, given historical steady frames. Our network is composed of a generative network with spatial transformer networks embedded in different layers, and generates a stable frame for the incoming unstable frame by computing an appropriate affine transformation. We also introduce an adversarial network to determine the stability of apiece of video. The network is trained directly using the pair of steady and unsteady videos. Experiments show that our method can produce similar results as traditional methods, moreover, it is capable of handling challenging unsteady video of low quality, where traditional methods fail, such as video with heavy noise or multiple exposures. Our method runs in real time, which is much faster than traditional methods. Sen-Zhe Xu 0001, Miao Wang 0004, Tai-Jiang Mu, Shi-Min Hu 0001 |
Comput. Graph. Forum | 4 |
| 2018 | Knowledge graph construction with structure and parameter learning for indoor scene designabstractWe consider the problem of learning a representation of both spatial relations and dependencies between objects for indoor scene design. We propose a novel knowledge graph framework based on the entity-relation model for representation of facts in indoor scene design, and further develop a weaklysupervised algorithm for extracting the knowledge graph representation from a small dataset using both structure and parameter learning. The proposed framework is flexible, transferable, and readable. We present a variety of computer-aided indoor scene design applications using this representation, to show the usefulness and robustness of the proposed framework. Song-Hai Zhang, Yukun Lai, Tai-Jiang Mu |
Comput. Vis. Media | 5 |
| 2017 | Image-based clothes changing systemabstractCurrent image-editing tools do not match up to the demands of personalized image manipulation, one application of which is changing clothes in usercaptured images. Previous work can change single color clothes using parametric human warping methods. In this paper, we propose an image-based clothes changing system, exploiting body factor extraction and content-aware image warping. Image segmentation and mask generation are first applied to the user input. Afterwards, we determine joint positions via a neural network. Then, body shape matching is performed and the shape of the model is warped to the user’s shape. Finally, head swapping is performed to produce realistic virtual results. We also provide a supervision and labeling tool for refinement and further assistance when creating a dataset. Zhao-Heng Zheng, Hao-Tian Zhang, Tai-Jiang Mu |
Comput. Vis. Media | 4 |
| 2015 | A response time model for abrupt changes in binocular disparity
Tai-Jiang Mu, Jia-Jia Sun, Ralph R. Martin, Shi-Min Hu 0001 |
Vis. Comput. | 1 |
| 2014 | Stereoscopic image completion and depth recovery
Tai-Jiang Mu, Ju-Hong Wang, Song-Pei Du, Shi-Min Hu 0001 |
Vis. Comput. | 1 |