EDBT 2026 Demo / reviewers in the wild / expert
Zhanyu Wang
dblp:123/2874
· DBLP profile ↗
24ranked-venue papers
9as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 10 · 6 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 7 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ReFINE: A Reward-Based Framework for Interpretable and Nuanced Evaluation of Radiology Report GenerationabstractAutomated radiology report generation (R2Gen) has advanced significantly, yet evaluation remains challenging due to the complexity of assessing report quality. Traditional metrics often misalign with human judgments, failing to identify specific deficiencies. To address this, we introduce ReFINE, a framework for training an Evaluation Model using a novel margin-based reward enforcement loss. This approach decomposes report quality into fine-grained sub-scores across user-defined criteria, improving interpretability. Leveraging GPT-4, we generate diverse training data with paired accepted and rejected reports to train our model under a reward-based system. The trained ReFINE Score provides both granular sub-scores and an aggregated quality assessment, enabling criterion-specific evaluation. Experimental results demonstrate ReFINE's superior alignment with human judgments, outperforming traditional metrics in model selection. Its robustness is validated across three expert-annotated datasets—including chest X-rays and multimodal reports covering 9 imaging modalities—and under two distinct scoring systems. Yunyi Liu, Yingshu Li 0001, Zhanyu Wang, Lingqiao Liu, Lei Wang 0001, Luping Zhou |
AAAI | 3 |
| 2025 | Retrieval-Augmented Multi-Modal Chain-of-Thoughts Reasoning for Large Language ModelsabstractThe advancement of Large Language Models (LLMs) has brought substantial attention to the Chain of Thought (CoT) approach, primarily due to its ability to enhance the capability of LLMs on complex reasoning tasks. Moreover, the CoT approach extends to multi-modal tasks of LLMs. However, the selection of optimal CoT demonstration examples for LLMs in multi-modal reasoning remains less explored due to the inherent complexity of multi-modal examples. In this paper, we introduce a novel approach that addresses this challenge by using retrieval mechanisms to dynamically select demonstration examples based on cross-modal and intra-modal similarities. Furthermore, we employ a Group Selection method to select examples containing rationales from different retrieval directions to promote the diversity of demonstration examples. To the best of our knowledge, we are the first to apply Retrieval-Augmented Generation (RAG) with CoT to complex multi-modal reasoning tasks. Through a series of experiments on two popular benchmarks, ScienceQA and MathVista, we demonstrate that our approach significantly improves the performance of GPT-4 by 6% on ScienceQA and 12.9% on MathVista. Additionally, it enhances the performance of GPT-4V on these two datasets by 2.7%, respectively, further advancing the capabilities of the most advanced LLMs and Large Multimodal Models (LMMs) for complex multi-modal reasoning tasks. Bingshuai Liu, Chenyang Lyu, Zijun Min, Zhanyu Wang, Jinsong Su, Longyue Wang |
IJCNN | 4 |
| 2025 | On the Taxonomy, Tasks, and Open-Challenges for Multimodal Large Language ModelsabstractIn recent years, the field of Artificial Intelligence has witnessed the emergence of Multimodal Large Language Models (MLLMs) that have significantly advanced the state-of-the-art in understanding and generating content across various data modalities. These models, capable of processing and integrating information from text, images, audio, and video, have opened new avenues for research and applications. Distinguished by their ability to understand and generation information with diverse modalities, such as text, image, audio and many others, MLLMs mark a significant step towards the final aim of Artificial General Intelligence (AGI). This comprehensive survey provides an in-depth examination of MLLMs, highlighting their evolutionary trajectory, current state-of-the-art developments, and prospective future directions. Specifically, we show taxonomy of MLLMs by their modalities to be processed and model architecture for aligning multiple modalities. Besides, we also present discussion regarding the different types of tasks related to MLLMs. The paper further delves into the pressing challenges confronted in this domain, such as data scarcity, computational complexity, ethical dilemmas, and privacy considerations. We analyze these issues in the context of both development and deployment of MLLMs. The survey comprehensively demonstrate and summarise the recent advances of the transformative influence of MLLMs while acknowledging their potential limitations, thereby outlining a prospective roadmap for future research endeavors in this rapidly developing field. Lecheng Yan, Jiahui Geng, Minghao Wu, Zhanyu Wang, Wenxi Li, Tianbo Ji, Shaochen Jiang, Chenyang Lyu |
SMC | 6 |
| 2025 | Differentially Private Bootstrap: New Privacy Analysis and Inference StrategiesabstractDifferentially private (DP) mechanisms protect individual-level information by introducing randomness into the statistical analysis procedure. Despite the availability of numerous DP tools, there remains a lack of general techniques for conducting statistical inference under DP. We examine a DP bootstrap procedure that releases multiple private bootstrap estimates to infer the sampling distribution and construct confidence intervals (CIs). Our privacy analysis presents new results on the privacy cost of a single DP bootstrap estimate, applicable to any DP mechanism, and identifies some misapplications of the bootstrap in the existing literature. For the composition of the DP bootstrap, we present a numerical method to compute the exact privacy cost of releasing multiple DP bootstrap estimates, and using the Gaussian-DP (GDP) framework (Dong et al., 2022) we show that the release of $B$ DP bootstrap estimates from mechanisms satisfying $(\mu/\sqrt{(2-2/\mathrm{e})B})$-GDP asymptotically satisfies $\mu$-GDP as $B$ goes to infinity. Then, we perform private statistical inference by post-processing the DP bootstrap estimates. We prove that our point estimates are consistent, our standard CIs are asymptotically valid, and both enjoy optimal convergence rates. To further improve the finite performance, we use deconvolution with DP bootstrap estimates to accurately infer the sampling distribution. We derive CIs for tasks such as population mean estimation, logistic regression, and quantile regression, and we compare them to existing methods using simulations and real-world experiments on 2016 Canada Census data. Our private CIs achieve the nominal coverage level and offer the first approach to private inference for quantile regression. Zhanyu Wang, Guang Cheng 0003, Jordan Awan |
J. Mach. Learn. Res. | 1 |
| 2025 | Can large language models effectively process and execute financial trading instructions?abstractThe development of large language models (LLMs) has created transformative opportunities for the financial industry, especially in the area of financial trading. However, how to integrate LLMs with trading systems has become a challenge. To address this problem, we propose an intelligent trade order recognition pipeline that enables the conversion of trade orders into a standard format for trade execution. The system improves the ability of human traders to interact with trading platforms while addressing the problem of misinformation acquisition in trade execution. In addition, we create a trade order dataset of 500 pieces of data to simulate the real-world trading scenarios. Moreover, we design several metrics to provide a comprehensive assessment of dataset reliability and the generative power of big models in finance by using five state-of-the-art LLMs on our dataset. The results show that most models generate syntactically valid JavaScript object notation (JSON) at high rates (about 80%–99%) and initiate clarifying questions in nearly all incomplete cases (about 90%–100%). However, end-to-end accuracy remains low (about 6%–14%), and missing information is substantial (about 12%–66%). Models also tend to over-interrogate—roughly 70%–80% of follow-ups are unnecessary—raising interaction costs and potential information-exposure risk. The research also demonstrates the feasibility of integrating our pipeline with the real-world trading systems, paving the way for practical deployment of LLM-based trade automation solutions. Yuda Wang, Zhanyu Wang, Mingwen Liu |
Frontiers Inf. Technol. Electron. Eng. | 5 |
| 2025 | Diagnostic Captioning by Cooperative Task Interactions and Sample-Graph ConsistencyabstractRadiographic images are similar to each other, making it challenging for diagnostic captioning to narrate fine-grained visual differences of clinical importance. In this paper, we propose a self-boosting framework integrating two novel strategies to learn tightly correlated image and text features for diagnostic captioning. The first strategy explicitly aligns image and text features through training an auxiliary task of image-text matching (ITM) jointly with the main task of report generation (RG) as two branches of a network model. The ITM branch explicitly learns image-text alignment and provides highly correlated visual and textual features for the RG branch to generate high-quality reports. The high-quality reports generated by RG branch, in turn, are utilized as additional harder negative samples to push the ITM branch to evolve towards better image-text alignment. These two branches help improve each other progressively, so that the whole model is self-boosted without requiring external resources. The second strategy aligns image-sample space and report-sample space to achieve consistent image and text feature embeddings. To achieve this, the sample graph of the embedded ground-truth reports is built and used as the target to train the sample graph of the embedded images so that the fine discrepancy in the ground-truth reports could be captured by the learned visual feature embeddings. Our proposed framework demonstrates its superiority on two medical report generation benchmarks, including the largest dataset MIMIC-CXR. Zhanyu Wang, Lei Wang 0001, Xiu Li 0001, Luping Zhou |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Enhancing Radiology Report Generation via Multi-Phased SupervisionabstractRadiology report generation using large language models has recently produced reports with more realistic styles and better language fluency. However, their clinical accuracy remains inadequate. Considering the significant imbalance between clinical phrases and general descriptions in a report, we argue that using an entire report for supervision is problematic as it fails to emphasize the crucial clinical phrases, which require focused learning. To address this issue, we propose a multi-phased supervision method, inspired by the spirit of curriculum learning where models are trained by gradually increasing task complexity. Our approach organizes the learning process into structured phases at different levels of semantical granularity, each building on the previous one to enhance the model. During the first phase, disease labels are used to supervise the model, equipping it with the ability to identify underlying diseases. The second phase progresses to use entity-relation triples to guide the model to describe associated clinical findings. Finally, in the third phase, we introduce conventional whole-report-based supervision to quickly adapt the model for report generation. Throughout the phased training, the model remains the same and consistently operates in the generation mode. As experimentally demonstrated, this proposed change in the way of supervision enhances report generation, achieving state-of-the-art performance in both language fluency and clinical accuracy. Our work underscores the importance of training process design in radiology report generation. Our code is available on https://github.com/zailongchen/MultiP-R2Gen. Zailong Chen, Yingshu Li 0002, Zhanyu Wang, Johan Barthelemy, Luping Zhou, Lei Wang 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2024 | KARGEN: Knowledge-Enhanced Automated Radiology Report Generation Using Large Language Models
Yingshu Li 0002, Zhanyu Wang, Yunyi Liu, Lei Wang 0001, Lingqiao Liu, Luping Zhou |
MICCAI (5) | 2 |
| 2024 | MRScore: Evaluating Medical Report with LLM-Based Reward System
Yunyi Liu, Zhanyu Wang, Yingshu Li 0002, Lingqiao Liu, Lei Wang 0001, Luping Zhou |
MICCAI (3) | 2 |
| 2024 | GPT4Video: A Unified Multimodal Large Language Model for lnstruction-Followed Understanding and Safety-Aware Generation
Zhanyu Wang, Longyue Wang, Zhen Zhao 0001, Minghao Wu, Chenyang Lyu, Deng Cai 0002, Luping Zhou, Shuming Shi 0001, Zhaopeng Tu |
ACM Multimedia | 1 |
| 2023 | METransformer: Radiology Report Generation by Transformer with Multiple Learnable Expert TokensabstractIn clinical scenarios, multi-specialist consultation could significantly benefit the diagnosis, especially for intricate cases. This inspires us to explore a “multi-expert joint diagnosis” mechanism to upgrade the existing “single expert” framework commonly seen in the current literature. To this end, we propose METransformer, a method to realize this idea with a transformer-based backbone. The key design of our method is the introduction of multiple learnable “expert” tokens into both the transformer encoder and decoder. In the encoder, each expert token interacts with both vision tokens and other expert tokens to learn to attend different image regions for image representation. These expert tokens are encouraged to capture complementary information by an orthogonal loss that minimizes their over-lap. In the decoder, each attended expert token guides the cross-attention between input words and visual tokens, thus influencing the generated report. A metrics-based expert voting strategy is further developed to generate the final report. By the multi-experts concept, our model enjoys the merits of an ensemble-based approach but through a manner that is computationally more efficient and supports more sophisticated interactions among experts. Experimental results demonstrate the promising performance of our proposed model on two widely used benchmarks. Last but not least, the framework-level innovation makes our work ready to incorporate advances on existing “single-expert” models to further improve its performance. Zhanyu Wang, Lingqiao Liu, Lei Wang 0001, Luping Zhou |
CVPR | 1 |
| 2023 | Stay in Grid: Improving Video Captioning via Fully Grid-Level RepresentationabstractVideo captioning is a challenging task of automatically generating natural and meaningful textual descriptions given some context videos. The state-of-the-art methods aggregate the spatial-wise information in the video encoder at the early stage, which has two drawbacks: 1) Early aggregation in the encoder can cause considerable spatial details missing, which may consequently lead to incorrect word choices in the following text encoder. 2) The spatial attention learned in the video encoder may not be compelling enough without text guidance. To solve these problems, we propose a Stay-in-Grid video CAPtioning method SGCAP, which makes full use of the grid-level spatial features and consists of a Bilinear Sequential Attention Encoder (BSAE) and a Cross-modal Sequential Attention Decoder (CSAD). The former explores and retains fully grid-level discriminative representations in the video encoder, while the latter performs the late spatial aggregation in the decoder to attend to the most relevant regions with the supervision of the input words. Experimental results demonstrate the effectiveness of our method on three public datasets, showing its superior performance over multiple state-of-the-art video captioning models. Source codes and the pre-trained models will be made available to the public. Mingkang Tang, Zhanyu Wang, Zhaoyang Zeng, Xiu Li 0001, Luping Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | MICO: Selective Search with Mutual Information Co-trainingabstractIn contrast to traditional exhaustive search, selective search first clusters documents into several groups before all the documents are searched exhaustively by a query, to limit the search executed within one group or only a few groups. Selective search is designed to reduce the latency and computation in modern large-scale search systems. In this study, we propose MICO, a Mutual Information CO-training framework for selective search with minimal supervision using the search logs. After training, MICO does not only cluster the documents, but also routes unseen queries to the relevant clusters for efficient retrieval. In our empirical experiments, MICO significantly improves the performance on multiple metrics of selective search and outperforms a number of existing competitive baselines. Zhanyu Wang, Hyokun Yun, Choon Hui Teo, Trishul Chilimbi |
COLING | 1 |
| 2022 | A Medical Semantic-Assisted Transformer for Radiographic Report Generation
Zhanyu Wang, Mingkang Tang, Lei Wang 0001, Xiu Li 0001, Luping Zhou |
MICCAI (3) | 1 |
| 2022 | Variance reduction on general adaptive stochastic mirror descent
Zhanyu Wang, Guang Cheng 0003 |
Mach. Learn. | 2 |
| 2022 | Automated Radiographic Report Generation Purely on Transformer: A Multicriteria Supervised ApproachabstractAutomated radiographic report generation is challenging in at least two aspects. First, medical images are very similar to each other and the visual differences of clinic importance are often fine-grained. Second, the disease-related words may be submerged by many similar sentences describing the common content of the images, causing the abnormal to be misinterpreted as the normal in the worst case. To tackle these challenges, this paper proposes a pure transformer-based framework to jointly enforce better visual-textual alignment, multi-label diagnostic classification, and word importance weighting, to facilitate report generation. To the best of our knowledge, this is the first pure transformer-based framework for medical report generation, which enjoys the capacity of transformer in learning long range dependencies for both image regions and sentence words. Specifically, for the first challenge, we design a novel mechanism to embed an auxiliary image-text matching objective into the transformer's encoder-decoder structure, so that better correlated image and text features could be learned to help a report to discriminate similar images. For the second challenge, we integrate an additional multi-label classification task into our framework to guide the model in making correct diagnostic predictions. Also, a term-weighting scheme is proposed to reflect the importance of words for training so that our model would not miss key discriminative information. Our work achieves promising performance over the state-of-the-arts on two benchmark datasets, including the largest dataset MIMIC-CXR. Zhanyu Wang, Lei Wang 0001, Xiu Li 0001, Luping Zhou |
IEEE Trans. Medical Imaging | 1 |
| 2021 | The Sample Complexity of Meta Sparse RegressionabstractThis paper addresses the meta-learning problem in sparse linear regression with infinite tasks. We assume that the learner can access several similar tasks. The goal of the learner is to transfer knowledge from the prior tasks to a similar but novel task. For $p$ parameters, size of the support set $k$, and $l$ samples per task, we show that $T \in O((k \log (p-k)) / l)$ tasks are sufficient in order to recover the common support of all tasks. With the recovered support, we can greatly reduce the sample complexity for estimating the parameter of the novel task, i.e., $l \in O(1)$ with respect to $T$ and $p$. We also prove that our rates are minimax optimal. A key difference between meta-learning and the classical multi-task learning, is that meta-learning focuses only on the recovery of the parameters of the novel task, while multi-task learning estimates the parameter of all tasks, which requires $l$ to grow with $T$. Instead, our efficient meta-learning estimator allows for $l$ to be constant with respect to $T$ (i.e., few-shot learning). Zhanyu Wang, Jean Honorio |
AISTATS | 1 |
| 2021 | A Self-Boosting Framework for Automated Radiographic Report GenerationabstractAutomated radiographic report generation is a challenging task since it requires to generate paragraphs describing fine-grained visual differences of cases, especially for those between the diseased and the healthy. Existing image captioning methods commonly target at generic images, and lack mechanism to meet this requirement. To bridge this gap, in this paper, we propose a self-boosting framework that improves radiographic report generation based on the cooperation of the main task of report generation and an auxiliary task of image-text matching. The two tasks are built as the two branches of a network model and influence each other in a cooperative way. On one hand, the image-text matching branch helps to learn highly text-correlated visual features for the report generation branch to output high quality reports. On the other hand, the improved reports produced by the report generation branch provide additional harder samples for the image-text matching branch and enforce the latter to improve itself by learning better visual and text feature representations. This, in turn, helps improve the report generation branch again. These two branches are jointly trained to help improve each other iteratively and progressively, so that the whole model is self-boosted without requiring external resources. Experimental results demonstrate the effectiveness of our method on two public datasets, showing its superior performance over multiple state-of-the-art image captioning and medical report generation methods. Zhanyu Wang, Luping Zhou, Lei Wang 0001, Xiu Li 0001 |
CVPR | 1 |
| 2021 | CLIP4Caption: CLIP for Video CaptionabstractVideo captioning is a challenging task since it requires generating sentences describing various diverse and complex videos. Existing video captioning models lack adequate visual representation due to the neglect of the existence of gaps between videos and texts. To bridge this gap, in this paper, we propose a CLIP4Caption framework that improves video captioning based on a CLIP-enhanced video-text matching network (VTM). This framework is taking full advantage of the information from both vision and language and enforcing the model to learn strongly text-correlated video features for text generation. Besides, unlike most existing models using LSTM or GRU as the sentence decoder, we adopt a Transformer structured decoder network to effectively learn the long-range visual and language dependency. Additionally, we introduce a novel ensemble strategy for captioning tasks. Experimental results demonstrate the effectiveness of our method on two datasets: 1) on MSR-VTT dataset, our method achieved a new state-of-the-art result with a significant gain of up to 10% in CIDEr; 2) on the private test data, our method ranking 2nd place in the ACM MM multimedia grand challenge 2021: Pre-training for Video Understanding Challenge. It is noted that our model is only trained on the MSR-VTT dataset. Mingkang Tang, Zhanyu Wang, Fengyun Rao, Xiu Li 0001 |
ACM Multimedia | 2 |
| 2020 | Directional Pruning of Deep Neural NetworksabstractIn the light of the fact that the stochastic gradient descent (SGD) often finds a flat minimum valley in the training loss, we propose a novel directional pruning method which searches for a sparse minimizer in or close to that flat region. The proposed pruning method does not require retraining or the expert knowledge on the sparsity level. To overcome the computational formidability of estimating the flat directions, we propose to use a carefully tuned $\ell_1$ proximal gradient algorithm which can provably achieve the directional pruning with a small learning rate after sufficient training. The empirical results demonstrate the promising results of our solution in highly sparse regime (92% sparsity) among many existing pruning methods on the ResNet50 with the ImageNet, while using only a slightly higher wall time and memory footprint than the SGD. Using the VGG16 and the wide ResNet 28x10 on the CIFAR-10 and CIFAR-100, we demonstrate that our solution reaches the same minima valley as the SGD, and the minima found by our solution and the SGD do not deviate in directions that impact the training loss. The code that reproduces the results of this paper is available at https://github.com/donlan2710/gRDA-Optimizer/tree/master/directional_pruning. Shih-Kang Chao, Zhanyu Wang, Yue Xing 0002, Guang Cheng 0003 |
NeurIPS | 2 |
| 2018 | BAUM: improving genome assembly by adaptive unique mapping and local overlap-layout-consensus approachabstractMotivation: It is highly desirable to assemble genomes of high continuity and consistency at low cost. The current bottleneck of draft genome continuity using the second generation sequencing (SGS) reads is primarily caused by uncertainty among repetitive sequences. Even though the single-molecule real-time sequencing technology is very promising to overcome the uncertainty issue, its relatively high cost and error rate add burden on budget or computation. Many long-read assemblers take the overlap-layout-consensus (OLC) paradigm, which is less sensitive to sequencing errors, heterozygosity and variability of coverage. However, current assemblers of SGS data do not sufficiently take advantage of the OLC approach. Results: Aiming at minimizing uncertainty, the proposed method BAUM, breaks the whole genome into regions by adaptive unique mapping; then the local OLC is used to assemble each region in parallel. BAUM can (i) perform reference-assisted assembly based on the genome of a close species (ii) or improve the results of existing assemblies that are obtained based on short or long sequencing reads. The tests on two eukaryote genomes, a wild rice Oryza longistaminata and a parrot Melopsittacus undulatus, show that BAUM achieved substantial improvement on genome size and continuity. Besides, BAUM reconstructed a considerable amount of repetitive regions that failed to be assembled by existing short read assemblers. We also propose statistical approaches to control the uncertainty in different steps of BAUM. Availability and implementation: http://www.zhanyuwang.xin/wordpress/index.php/2017/07/21/baum. Supplementary information: Supplementary data are available at Bioinformatics online. Anqi Wang 0002, Zhanyu Wang, Lei M. Li |
Bioinform. | 2 |
| 2014 | Demonstration abstract: upper body motion capture system using inertial sensors
Jian Wu 0016, Zhanyu Wang, Suraj Raghuraman, B. Prabhakaran 0001, Roozbeh Jafari |
IPSN | 2 |
| 2013 | A 3D tele-immersion streaming approach using skeleton-based predictionabstract3D collaborative Tele-Immersive environments allow reconstruction of real world 3D scenes in the virtual world across multiple physical locations. This kind of reconstruction results in a lot of 3D data being transmitted over the internet in real time. The current systems allow for transmission at low frame rates due to the large volume of data and network bandwidth restrictions. In this paper we propose a prediction based approach that generates future frames by animating the live model based on few skeleton points. By doing so the magnitude of data transmitted is reduced to few hundred bytes. The prediction errors are corrected when an entire frame is received. This approach allows minimal amounts (few bytes) of data to be transmitted per frame, thus allowing for high frame rates and still maintain an acceptable visual quality of reconstruction at the receiver side. Suraj Raghuraman, Karthik Venkatraman, Zhanyu Wang, B. Prabhakaran 0001, Xiaohu Guo |
ACM Multimedia | 3 |
| 2012 | Immersive multiplayer tennis with microsoft kinect and body sensor networksabstractWe present an immersive gaming demonstration using the minimum amount of wearable sensors. The game demonstrated is two-player tennis. We combine a virtual environment with real 3D representations of physical objects like the players and the tennis racquet (if available). The main objective of the game is to provide as real an experience of tennis as possible, while also being as less intrusive as possible. The game is played across a network, and this opens the possibility of two remote players playing a game together on a single virtual tennis pitch. The Microsoft Kinect sensors are used to obtain a 3D point cloud and a skeletal map representation of the player. This 3D point cloud is mapped on to the virtual tennis pitch. We also use a wireless wearable Attitude and Heading Reference System (AHRS) mote, which is strapped onto the wrist of the players. This mote gives us precise information about the movement (swing, rotation etc.) of the playing arm. This information along with the skeletal map is used to implement the physics of the game. Using this game we demonstrate our solutions for simultaneous data acquisition, 3D point-cloud mapping in a virtual space, use of the Kinect and AHRS sensors to calibrate real and virtual objects and for interaction of virtual objects with a 3D point cloud. Suraj Raghuraman, Karthik Venkatraman, Zhanyu Wang, Jian Wu 0016, Jacob Clements, Reza Lotfian, B. Prabhakaran 0001, Xiaohu Guo, Roozbeh Jafari, Klara Nahrstedt |
ACM Multimedia | 3 |