VLDB 2026 Research / reviewers in the wild / expert
Wen-Cheng Chen
dblp:32/10138
· DBLP profile ↗
13ranked-venue papers
4as first author
7since 2021 · last 2023
0000-0002-7492-3591ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Scalable Spatial Memory for Scene Rendering and NavigationabstractNeural scene representation and rendering methods have shown promise in learning the implicit form of scene structure without supervision. However, the implicit representation learned in most existing methods is non-expandable and cannot be inferred online for novel scenes, which makes the learned representation difficult to be applied across different reinforcement learning (RL) tasks. In this work, we introduce Scene Memory Network (SMN) to achieve online spatial memory construction and expansion for view rendering in novel scenes. SMN models the camera projection and back-projection as spatially aware memory control processes, where the memory values store the information of the partial 3D area, and the memory keys indicate the position of that area. The memory controller can learn the geometry property from observations without the camera's intrinsic parameters and depth supervision. We further apply the memory constructed by SMN to exploration and navigation tasks. The experimental results reveal the generalization ability of our proposed SMN in large-scale scene synthesis and its potential to improve the performance of spatial RL tasks. Wen-Cheng Chen, Chu-Song Chen, Walon Wei-Chen Chiu, Min-Chun Hu 0001 |
AAAI | 1 |
| 2023 | Offensive Tactics Recognition in Broadcast Basketball Videos Based on 2D Camera View Player HeatmapsabstractIt is essential for sports teams to review their offensive and defensive tactical execution performance as well as understand their opponents’ tactics in order to identify effective counterattack strategies. This study focuses on basketball offensive tactics recognition based on 2D camera view heatmaps. Most of the current tactics recognition methods learn the spatiotemporal correlation of players based on top-view trajectory information. To obtain correct top-view player trajectories, robust camera calibration and player tracking techniques are indispensable. However, for broadcast videos having large camera movement, serious player occlusions, and similar players’ jerseys, it is quite challenging to obtain accurate camera parameters and player tracking results, resulting in poor tactical analysis performance. Instead of applying camera calibration and player tracking, this study attempts to design a tactics recognition method that directly predicts the tactics class from 2D camera-view player heatmaps in the inference phase. Our proposed method uses a recurrent convolutional neural network with coordinate embedding to directly identify the tactics. Moreover, an auxiliary top-view player trajectory reconstruction module is added in the training phase to acquire better latent codes to represent the tactics. The experimental results show that for both supervised and unsupervised settings, our proposed method achieves comparable accuracy to the current tactics classification methods that rely on perfect top-view trajectory input. subst Nico, Tse-Yu Pan, Herman Prawiro, Jain-Wei Peng, Wen-Cheng Chen, Hung-Kuo Chu, Min-Chun Hu 0001 |
ICMR | 5 |
| 2023 | Domain Invariant Vision Transformer Learning for Face Anti-spoofingabstractExisting face anti-spoofing (FAS) models have achieved high performance on specific datasets. However, for the application of real-world systems, the FAS model should generalize to the data from unknown domains rather than only achieve good results on a single baseline. As vision transformer models have demonstrated astonishing performance and strong capability in learning discriminative information, we investigate applying transformers to distinguish the face presentation attacks over unknown domains. In this work, we propose the Domain-invariant Vision Transformer (DiVT) for FAS, which adopts two losses to improve the generalizability of the vision transformer. First, a concentration loss is employed to learn a domain-invariant representation that aggregates the features of real face data. Second, a separation loss is utilized to union each type of attack from different domains. The experimental results show that our proposed method achieves state-of-the-art performance on the protocols of domain-generalized FAS tasks. Compared to previous domain generalization FAS models, our proposed method is simpler but more effective. Chen-Hao Liao, Wen-Cheng Chen, Hsuan-Tung Liu, Yi-Ren Yeh, Min-Chun Hu 0001, Chu-Song Chen |
WACV | 2 |
| 2022 | Learning Robust Latent Space of Basketball Player Trajectories for Tactics AnalysisabstractTactic analysis of trajectory data plays an important role in the field of sports science. However, the tactical labels of trajectories are usually insufficient, which limits the deep models to learn the general concept of tactics. In this work, we combine the recurrent variational autoencoder with attention module to learn a robust latent space of the basketball offensive trajectories without the need of labeled data. The learned latent space can be further applied to advanced tasks such as supervised tactics classification, unsupervised tactics clustering, and defensive trajectory generation. The experimental results show that our proposed model achieves better or competitive performance than the previous methods. Furthermore, our proposed model is the first one that focuses on zero-shot tactics clustering and the performance even outperforms the previous methods with supervised and unsupervised settings, which shows the promise of tactics analysis for trajectories in the case of limited labeled data. Yang-Sheng Chao, Wen-Cheng Chen, Jain-Wei Peng, Min-Chun Hu 0001 |
ICME | 2 |
| 2022 | Instant Basketball Defensive Trajectory GenerationabstractTactic learning in virtual reality (VR) has been proven to be effective for basketball training. Endowed with the ability of generating virtual defenders in real time according to the movement of virtual offenders controlled by the user, a VR basketball training system can bring more immersive and realistic experiences for the trainee. In this article, an autoregressive generative model for instantly producing basketball defensive trajectory is introduced. We further focus on the issue of preserving the diversity of the generated trajectories. A differentiable sampling mechanism is adopted to learn the continuous Gaussian distribution of player position. Moreover, several heuristic loss functions based on the domain knowledge of basketball are designed to make the generated trajectories assemble real situations in basketball games. We compare the proposed method with the state-of-the-art works in terms of both objective and subjective manners. The objective manner compares the average position, velocity, and acceleration of the generated defensive trajectories with the real ones to evaluate the fidelity of the results. In addition, more high-level aspects such as the empty space for offender and the defensive pressure of the generated trajectory are also considered in the objective evaluation. As for the subjective manner, visual comparison questionnaires on the proposed and other methods are thoroughly conducted. The experimental results show that the proposed method can achieve better performance than previous basketball defensive trajectory generation works in terms of different evaluation metrics. Wen-Cheng Chen, Wan-Lun Tsai, Huan-Hua Chang, Min-Chun Hu 0001, Wei-Ta Chu |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2021 | STR-GQN: Scene Representation and Rendering for Unknown Cameras Based on Spatial Transformation RoutingabstractGeometry-aware modules are widely applied in recent deep learning architectures for scene representation and rendering. However, these modules require intrinsic camera information that might not be obtained accurately. In this paper, we propose a Spatial Transformation Routing (STR) mechanism to model the spatial properties without applying any geometric prior. The STR mechanism treats the spatial transformation as the message passing process, and the relation between the view poses and the routing weights is modeled by an end-to-end trainable neural network. Besides, an Occupancy Concept Mapping (OCM) framework is proposed to provide explainable rationals for scene-fusion processes. We conducted experiments on several datasets and show that the proposed STR mechanism improves the performance of the Generative Query Network (GQN). The visualization results reveal that the routing process can pass the observed information from one location of some view to the associated location in the other view, which demonstrates the advantage of the proposed model in terms of spatial cognition. Wen-Cheng Chen, Min-Chun Hu 0001, Chu-Song Chen |
ICCV | 1 |
| 2021 | Semi-supervised Many-to-many Music Timbre TransferabstractThis work presents a music timbre transfer model that aims to transfer the style of a music clip while preserving the semantic content. Compared to the existing music timbre transfer models, our model can achieve many-to-many timbre transfer between different instruments. The proposed method is based an autoencoder framework, which comprises two pretrained encoders trained in a supervised manner and one decoder trained in an unsupervised manner. To learn more representative features for the encoders, we produced a parallel dataset, called MI-Para, which is synthesized from MIDI files and digital audio workstations (DAW). Both the objective and the subjective evaluation results showed the effectiveness of the proposed framework. To scale up the application scenario, we also demonstrate that our model can achieve style transfer by training in a semi-supervised manner with a smaller parallel dataset. Yu-Chen Chang, Wen-Cheng Chen, Min-Chun Hu 0001 |
ICMR | 2 |
| 2020 | An autoregressive generation model for producing instant basketball defensive trajectoryabstractLearning basketball tactic via virtual reality environment requires real-time feedback to improve the realism and interactivity. For example, the virtual defender should move immediately according to the player's movement. In this paper, we proposed an autoregressive generative model for basketball defensive trajectory generation. To learn the continuous Gaussian distribution of player position, we adopt a differentiable sampling process to sample the candidate location with a standard deviation loss, which can preserve the diversity of the trajectories. Furthermore, we design several additional loss functions based on the domain knowledge of basketball to make the generated trajectories match the real situation in basketball games. The experimental results show that the proposed method can achieve better performance than previous works in terms of different evaluation metrics. Huan-Hua Chang, Wen-Cheng Chen, Wan-Lun Tsai, Min-Chun Hu 0001, Wei-Ta Chu |
MMAsia | 2 |
| 2020 | Framework Design for Multiplayer Motion Sensing Game in Mixture Reality
Chih-Yao Chang, Bo-I Chuang, Chi-Chun Hsia, Wen-Cheng Chen, Min-Chun Hu 0001 |
MMM (2) | 4 |
| 2019 | IMMVP: An Efficient Daytime and Nighttime On-Road Object DetectorabstractIt is hard to detect on-road objects under various lighting conditions. To improve the quality of the classifier, three techniques are used. We define subclasses to separate daytime and nighttime samples. Then we skip similar samples in the training set to prevent overfitting. With the help of the outside training samples, the detection accuracy is also improved. To detect objects in an edge device, Nvidia Jetson TX2 platform, we exert the lightweight model ResNet-18 FPN as the backbone feature extractor. The FPN (Feature Pyramid Network) generates good features for detecting objects over various scales. With Cascade R-CNN technique, the bounding boxes are iteratively refined for better results. Cheng-En Wu, Yi-Ming Chan, Wen-Cheng Chen, Chu-Song Chen |
MMSP | 4 |
| 2018 | Syncgan: Synchronize the Latent Spaces of Cross-Modal Generative Adversarial NetworksabstractGenerative adversarial network (GAN) has achieved impressive success on cross-domain generation, but it faces difficulty in cross-modal generation due to the lack of a common distribution between heterogeneous data. Most existing methods of conditional based cross-modal GANs adopt the strategy of one-directional transfer and have achieved preliminary success on text-to-image transfer. Instead of learning the transfer between different modalities, we aim to learn a synchronous latent space representing the cross-modal common concept. A novel network component named synchronizer is proposed in this work to judge whether the paired data is synchronous/corresponding or not, which can constrain the latent space of generators in the GANs. Our GAN model, named as SyncGAN, can successfully generate synchronous data (e.g., a pair of image and sound) from identical random noise. For transforming data from one modality to another, we recover the latent code by inverting the mappings of a generator and use it to generate data of different modality. In addition, the proposed model can achieve semi-supervised learning, which makes our model more flexible for practical applications. Wen-Cheng Chen, Chien-Wen Chen, Min-Chun Hu 0001 |
ICME | 1 |
| 2012 | Quasi-sliding mode control for a class of multivariable discrete time bilinear systemsabstractIn this paper, we adopt the two quasi-sliding mode control for a class of multivariable discrete time bilinear systems. By applying the proposed control strategies, the system states would move into a boundary layer of the sliding surface and remain in the boundary layer in finite time. Thus, the stability of the overall system could be guaranteed. Finally, the computer simulation is applied to illustrate the control performance of the proposed control strategies. Chung-Chun Kung, Ti-Hung Chen, Wen-Cheng Chen, Jui-Yiao Su |
SMC | 3 |
| 2003 | A low power VLSI implementation for variable length decoder in MPEG-1 layer IIIabstractMPEG layer III (MP3) audio coding algorithm was a widely used audio coding standard. It involves several complex coding techniques and is therefore difficult to create an efficient architecture design. The variable length decoding was an important part, which needs great amount of search and memory read/write operations. In this paper a data driven variable length decoding algorithm is presented, which can exploit the signal statistics of variable length codes to reduce power. The decoder was designed based on simplicity and low-cost, low power consumption while retaining the high efficiency requirements. Tsung-Han Tsai 0001, Wen-Cheng Chen, Chun-Nan Liu |
ICME | 2 |