EDBT 2026 Demo / reviewers in the wild / expert
Tinghuai Wang
dblp:55/8166
· DBLP profile ↗
28ranked-venue papers
15as first author
5since 2021 · last 2025
0000-0002-7863-3516ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 11 first-author · 2 since 2021Artificial intelligence and machine learning · 15 · 8 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Reinforcement learning · 36% Video understanding and tracking · 28% Probabilistic and Bayesian machine learning · 17% | |
| Computer graphics and multimedia
2 papers |
Rendering · 46% Image and video processing · 32% Visual content generation and editing · 23% |
Topics — the 18 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
hierarchical reinforcement learning |
2.3 | 3 | 2025 | Hierarchical Reinforcement Learning with Uncertainty-Guided Diffusional Subgoals · ICML 2025 Probabilistic Subgoal Representations for Hierarchical Reinforcement Learning · ICML 2024 State-Conditioned Adversarial Subgoal Generation · AAAI 2023 |
Machine learning › Reinforcement learning › hierarchical reinforcement learning › option discovery
subgoal discovery |
1.5 | 2 | 2025 | Hierarchical Reinforcement Learning with Uncertainty-Guided Diffusional Subgoals · ICML 2025 State-Conditioned Adversarial Subgoal Generation · AAAI 2023 |
Machine learning › Generative modeling › diffusion model
conditional diffusion model |
0.9 | 1 | 2025 | Hierarchical Reinforcement Learning with Uncertainty-Guided Diffusional Subgoals · ICML 2025 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | Hierarchical Reinforcement Learning with Uncertainty-Guided Diffusional Subgoals · ICML 2025 |
Computer vision › Video understanding and tracking › video question answering
long-form video question answering |
0.9 | 1 | 2025 | ReWind: Understanding Long Videos with Instructed Learnable Memory · CVPR 2025 |
Computer vision › Video understanding and tracking
long video understanding |
0.9 | 1 | 2025 | ReWind: Understanding Long Videos with Instructed Learnable Memory · CVPR 2025 |
Computer vision › Video understanding and tracking
video question answering |
0.9 | 1 | 2025 | ReWind: Understanding Long Videos with Instructed Learnable Memory · CVPR 2025 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process |
0.8 | 1 | 2024 | Probabilistic Subgoal Representations for Hierarchical Reinforcement Learning · ICML 2024 |
Machine learning › Probabilistic and Bayesian machine learning
probabilistic representation |
0.8 | 1 | 2024 | Probabilistic Subgoal Representations for Hierarchical Reinforcement Learning · ICML 2024 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
graphical model inference |
0.3 | 1 | 2017 | Cross-Granularity Graph Inference for Semantic Video Object Segmentation · IJCAI 2017 |
Computer vision › Video understanding and tracking
video object segmentation |
0.3 | 1 | 2017 | Cross-Granularity Graph Inference for Semantic Video Object Segmentation · IJCAI 2017 |
Computer vision › Vision and language
temporal grounding |
0.3 | 1 | 2025 | ReWind: Understanding Long Videos with Instructed Learnable Memory · CVPR 2025 |
Visual content generation and editing › stylization
image stylization |
0.2 | 1 | 2013 | State of the "Art": A Taxonomy of Artistic Stylization Techniques for Images and Video · IEEE Trans. Vis. Comput. Graph. 2013 |
Rendering
non-photorealistic rendering |
0.2 | 1 | 2013 | State of the "Art": A Taxonomy of Artistic Stylization Techniques for Images and Video · IEEE Trans. Vis. Comput. Graph. 2013 |
Rendering › non-photorealistic rendering
painterly rendering |
0.2 | 1 | 2013 | State of the "Art": A Taxonomy of Artistic Stylization Techniques for Images and Video · IEEE Trans. Vis. Comput. Graph. 2013 |
Image and video processing
video segmentation |
0.1 | 1 | 2012 | Probabilistic Motion Diffusion of Labeling Priors for Coherent Video Segmentation · IEEE Trans. Multim. 2012 |
Image and video processing › image segmentation › graph-based segmentation
graph cut segmentation |
0.0 | 1 | 2012 | Probabilistic Motion Diffusion of Labeling Priors for Coherent Video Segmentation · IEEE Trans. Multim. 2012 |
Image and video processing › image segmentation
superpixel segmentation |
0.0 | 1 | 2012 | Probabilistic Motion Diffusion of Labeling Priors for Coherent Video Segmentation · IEEE Trans. Multim. 2012 |
Methods — techniques the papers use, named apart from their topics
gaussian process · 1.6vision-language model · 0.9learnable memory · 0.9large language model · 0.9diffusion model · 0.9cross-attention · 0.9adaptive frame selection · 0.9hierarchical reinforcement learning · 0.8discriminator network · 0.7adversarial learning · 0.7image gradient analysis · 0.2edge-aware filtering · 0.2multi-label graph cut · 0.1motion diffusion model · 0.1mean shift · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ReWind: Understanding Long Videos with Instructed Learnable MemoryabstractVision-Language Models (VLMs) are crucial for applications requiring integrated understanding textual and visual information. However, existing VLMs struggle with long videos due to computational inefficiency, memory limitations, and difficulties in maintaining coherent understanding across extended sequences. To address these challenges, we introduce ReWind, a novel memory-based VLM designed for efficient long video understanding while preserving temporal fidelity. ReWind operates in a two-stage framework. In the first stage, ReWind maintains a dynamic learnable memory module with a novel read-perceive-write cycle that stores and updates instruction-relevant visual information as the video unfolds. This module utilizes learnable queries and cross-attentions between memory contents and the input stream, ensuring low memory requirements by scaling linearly with the number of tokens. In the second stage, we propose an adaptive frame selection mechanism guided by the memory content to identify instruction-relevant key moments. It enriches the memory representations with detailed spatial information by selecting a few high-resolution frames, which are then combined with the memory contents and fed into a Large Language Model (LLM) to generate the final answer. We empirically demonstrate ReWind’s superior performance in visual question answering (VQA) and temporal grounding tasks, surpassing previous methods on long video benchmarks. Notably, ReWind achieves a +13% score gain and a +12% accuracy improvement on the MovieChat-1K VQA dataset and an +8% mIoU increase on Charades-STA for temporal grounding. Anxhelo Diko, Tinghuai Wang, Wassim Swaileh, Shiyan Sun, Ioannis Patras |
CVPR | 2 |
| 2025 | Hierarchical Reinforcement Learning with Uncertainty-Guided Diffusional SubgoalsabstractHierarchical reinforcement learning (HRL) learns to make decisions on multiple levels of temporal abstraction. A key challenge in HRL is that the low-level policy changes over time, making it difficult for the high-level policy to generate effective subgoals. To address this issue, the high-level policy must capture a complex subgoal distribution while also accounting for uncertainty in its estimates. We propose an approach that trains a conditional diffusion model regularized by a Gaussian Process (GP) prior to generate a complex variety of subgoals while leveraging principled GP uncertainty quantification. Building on this framework, we develop a strategy that selects subgoals from both the diffusion policy and GP’s predictive mean. Our approach outperforms prior HRL methods in both sample efficiency and performance on challenging continuous control benchmarks. Vivienne Huiling Wang, Tinghuai Wang, Joni Pajarinen |
ICML | 2 |
| 2024 | Probabilistic Subgoal Representations for Hierarchical Reinforcement LearningabstractIn goal-conditioned hierarchical reinforcement learning (HRL), a high-level policy specifies a subgoal for the low-level policy to reach. Effective HRL hinges on a suitable subgoal representation function, abstracting state space into latent subgoal space and inducing varied low-level behaviors. Existing methods adopt a subgoal representation that provides a deterministic mapping from state space to latent subgoal space. Instead, this paper utilizes Gaussian Processes (GPs) for the first probabilistic subgoal representation. Our method employs a GP prior on the latent subgoal space to learn a posterior distribution over the subgoal representation functions while exploiting the long-range correlation in the state space through learnable kernels. This enables an adaptive memory that integrates long-range subgoal information from prior planning steps allowing to cope with stochastic uncertainties. Furthermore, we propose a novel learning objective to facilitate the simultaneous learning of probabilistic subgoal representations and policies within a unified framework. In experiments, our approach outperforms state-of-the-art baselines in standard benchmarks but also in environments with stochastic elements and under diverse reward conditions. Additionally, our model shows promising capabilities in transferring low-level policies across different tasks. Vivienne Huiling Wang, Tinghuai Wang, Wenyan Yang, Joni-Kristian Kämäräinen, Joni Pajarinen |
ICML | 2 |
| 2023 | State-Conditioned Adversarial Subgoal GenerationabstractHierarchical reinforcement learning (HRL) proposes to solve difficult tasks by performing decision-making and control at successively higher levels of temporal abstraction. However, off-policy HRL often suffers from the problem of a non-stationary high-level policy since the low-level policy is constantly changing. In this paper, we propose a novel HRL approach for mitigating the non-stationarity by adversarially enforcing the high-level policy to generate subgoals compatible with the current instantiation of the low-level policy. In practice, the adversarial learning is implemented by training a simple state conditioned discriminator network concurrently with the high-level policy which determines the compatibility level of subgoals. Comparison to state-of-the-art algorithms shows that our approach improves both learning efficiency and performance in challenging continuous control tasks. Vivienne Huiling Wang, Joni Pajarinen, Tinghuai Wang, Joni-Kristian Kämäräinen |
AAAI | 3 |
| 2022 | Tradeoffs in the Spatial and Spectral Resolution of Airborne Hyperspectral Imaging Systems: A Crop Identification Case StudyabstractAirborne hyperspectral images are used for crop identification with a high classification accuracy because of their high spectral resolution, spatial resolution, and signal-to-noise ratio (SNR). However, the tradeoffs between the three core parameters of a hyperspectral imager (SNR, spatial resolution, and spectral resolution) should be considered for designing an efficient imaging system. Only a few reported studies on the analysis of the impact of SNR on identification accuracy are available. Further, the tradeoffs and mutual interactions among these parameters are rarely considered. In this empirical study, our aim was to understand the relationship among the core parameters and their effects on crop identification accuracy by analyzing the tradeoffs and mutual interactions among these parameters. We analyzed the hyperspectral images of a typical plain agricultural area in Xiongan, China, acquired by the newly developed sensor airborne multimodular imaging spectrometer (AMMIS). The fundamental images were transformed to form datasets with different ranges of spectral resolution, spatial resolution, and SNR using data reconstruction methods. We adopted the classification and regression tree (CART), random forest (RF), and k-nearest neighbor (kNN) classifiers, and observed the overall accuracy (OA) across the degraded hyperspectral datasets. The experimental results indicated that the OA decreased with a decreasing SNR. As the spectral resolution became coarser, the OA first increased, plateaued, and then decreased. However, the OA increased with decreasing spatial resolution. This study was performed with the goal of bridging the knowledge gap between the back-end hyperspectral sensor designing and its front-end applications. Jianxin Jia, Jinsong Chen 0001, Xiaorou Zheng, Yueming Wang 0002, Shanxin Guo, Haibin Sun 0002, Changhui Jiang, Mika Karjalainen, Kirsi Karila, Zhiyong Duan, Tinghuai Wang, Juha Hyyppä, Yuwei Chen 0005 |
IEEE Trans. Geosci. Remote. Sens. | 11 |
| 2019 | Simultaneously Learning Architectures and Features of Deep Neural Networks
Tinghuai Wang, Lixin Fan |
ICANN (2) | 1 |
| 2019 | Graph-Boosted Attentive Network for Semantic Body Parsing
Tinghuai Wang |
ICANN (3) | 1 |
| 2019 | Portrait Instance Segmentation for Mobile DevicesabstractAccurate and efficient portrait instance segmentation has become a crucial enabler for many multimedia applications on mobile devices. We present a novel convolutional neural network (CNN) architecture to explicitly address the long standing problems in portrait segmentation, i.e., semantic coherence and boundary localization. Specifically, we propose a cross-granularity categorical attention mechanism leveraging the deep supervisions to close the semantic gap of CNN feature hierarchy by imposing consistent category-oriented information across layers. Furthermore, a cross-granularity boundary enhancement module is proposed to boost the boundary awareness of deep layers by integrating the shape context cues from shallow layers of the network. We further propose a novel and efficient non-parametric affinity model to achieve efficient instance segmentation on mobile devices. We present a portrait image dataset with instance level annotations dedicated to evaluating portrait instance segmentation algorithms. We evaluate our approach on challenging datasets which obtains state-of-the-art results. Lingyu Zhu 0001, Tinghuai Wang, Emre Aksu, Joni-Kristian Kämäräinen |
ICME | 2 |
| 2018 | Non-parametric Contextual Relationship Learning for Semantic Video Object Segmentation
Tinghuai Wang |
CIARP | 1 |
| 2018 | Context Propagation from Proposals for Semantic Video Object SegmentationabstractIn this paper, we propose a novel approach to learning semantic contextual relationships in videos for semantic object segmentation. Our algorithm derives the semantic contexts from video object proposals which encode the key evolution of objects and the relationship among objects over the spatio-temporal domain. This semantic contexts are propagated across the video to estimate the pairwise contexts between all pairs of local superpixels which are integrated into a conditional random field in the form of pairwise potentials and infers the per-superpixel semantic labels. The experiments demonstrate that our contexts learning and propagation model effectively improves the robustness of resolving visual ambiguities in semantic video object segmentation compared with the state-of-the-art methods. Tinghuai Wang |
ICIP | 1 |
| 2017 | Submodular video object proposal selection for semantic object segmentationabstractLearning a data-driven spatio-temporal semantic representation of the objects is the key to coherent and consistent labelling in video. This paper proposes to achieve semantic video object segmentation by learning a data-driven representation which captures the synergy of multiple instances from continuous frames. To prune the noisy detections, we exploit the rich information among multiple instances and select the discriminative and representative subset. This selection process is formulated as a facility location problem solved by maximising a submodular function. Our method retrieves the longer term contextual dependencies which underpins a robust semantic video object segmentation algorithm. We present extensive experiments on a challenging dataset that demonstrate the superior performance of our approach compared with the state-of-the-art methods. Tinghuai Wang |
ICIP | 1 |
| 2017 | Cross-Granularity Graph Inference for Semantic Video Object SegmentationabstractWe address semantic video object segmentation via a novel cross-granularity hierarchical graphical model to integrate tracklet and object proposal reasoning with superpixel labeling. Tracklet characterizes varying spatial-temporal relations of video object which, however, quite often suffers from sporadic local outliers. In order to acquire high-quality tracklets, we propose a transductive inference model which is capable of calibrating short-range noisy object tracklets with respect to long-range dependencies and high-level context cues. In the center of this work lies a new paradigm of semantic video object segmentation beyond modeling appearance and motion of objects locally, where the semantic label is inferred by jointly exploiting multi-scale contextual information and spatial-temporal relations of video object. We evaluate our method on two popular semantic video object segmentation benchmarks and demonstrate that it advances the state-of-the-art by achieving superior accuracy performance than other leading methods. Tinghuai Wang, Ke Chen 0004, Joni-Kristian Kämäräinen |
IJCAI | 2 |
| 2016 | Semi-supervised Domain Adaptation for Weakly Labeled Semantic Video Object Segmentation
Tapani Raiko, Lasse Lensu, Tinghuai Wang, Juha Karhunen |
ACCV (1) | 4 |
| 2016 | Boosting objectness: Semi-supervised learning for object detection and segmentation in multi-view imagesabstractThis paper presents a method to detect and segment recurring object from multi-view images. Given a sequence of images of an object captured by multiple cameras, the method firstly detects sparse object-like regions utilizing generic region proposals. We propose a semi-supervised framework to exploit both appearance cues learned from rudimentary detections of object-like regions, and the intrinsic geometric structures within multi-view data. This framework generates a diverse set of object proposals in all views which underpins a robust object segmentation method to handle objects with complex shape and topologies, as well as scenarios where the object and background exhibit similar color distributions. Tinghuai Wang |
ICASSP | 2 |
| 2016 | Primary object discovery and segmentation in videos via graph-based transductive inference
Tinghuai Wang |
Comput. Vis. Image Underst. | 2 |
| 2015 | Robust interactive image segmentation with weak supervision for mobile touch screen devicesabstractIn this paper, we present a robust and efficient approach for segmenting images with less and intuitive user interaction, particularly targeted for mobile touch screen devices. Our approach combines geodesic distance information with the flexibility of level set methods in energy minimization, leveraging the complementary strengths of each to promote accurate boundary placement and strong region connectivity while requiring less user interaction. To maximize the user-provided prior knowledge, we further propose a weakly supervised seed generation algorithm which enables image object segmentation without user-provided background seeds. Our approach provides a practical solution for visual object cutout on mobile touch screen devices, facilitating various media manipulation applications. We describe such a use case to selectively create oil painting effects on images. We demonstrate that our approach is less sensitive to seed placement and better at edge localization, whilst requiring less user interaction, compared with the state-of-the-art methods. Tinghuai Wang, Lixin Fan |
ICME | 1 |
| 2015 | A weakly supervised geodesic level set framework for interactive image segmentation
Tinghuai Wang, Lixin Fan |
Neurocomputing | 1 |
| 2014 | Graph Transduction Learning of Object Proposals for Video Object Segmentation
Tinghuai Wang |
ACCV (4) | 1 |
| 2014 | Wide Baseline Multi-view Video Matting Using a Hybrid Markov Random FieldabstractWe describe a novel framework for segmenting a time- and view-coherent foreground matte sequence from synchronised multiple view video. We construct a Markov Random Field (MRF) comprising links between super pixels corresponded across views, and links between super pixels and their constituent pixels. Texture, colour and disparity cues are incorporated to model foreground appearance. We solve using a multi-resolution iterative approach enabling an eight view high definition (HD) frame to be processed in less than a minute. Furthermore we incorporate a temporal diffusion process introducing a prior on the MRF using information propagated from previous frames, and a facility for optional user correction. The result is a set of temporally coherent mattes solved for simultaneously across views for each frame, exploiting similarities across views and time. Tinghuai Wang, John P. Collomosse, Adrian Hilton 0001 |
ICPR | 1 |
| 2014 | TouchCut: Fast image and video segmentation using single-touch interaction
Tinghuai Wang, John P. Collomosse |
Comput. Vis. Image Underst. | 1 |
| 2013 | Learnable Stroke Models for Example-based Portrait PaintingabstractWe present a novel algorithm for stylizing photographs into portrait paintings comprised of curved brush strokes. Rather than drawing upon a prescribed set of heuristics to place strokes, our system learns a flexible model of artistic style by analyzing training data from a human artist. Given a training pair — a source image and painting of that image—a non-parametric model of style is learned by observing the geometry and tone of brush strokes local to image features. A Markov Random Field (MRF) enforces spatial coherence of style parameters. Style models local to facial features are learned using a semantic segmentation of the input face image, driven by a combination of an Active Shape Model and Graph-cut. We evaluate style transfer between a variety of training and test images, demonstrating a wide gamut of learned brush and shading styles. Tinghuai Wang, John P. Collomosse, Andrew Hunter, Darryl Greig |
BMVC | 1 |
| 2013 | Markov random fields for sketch based video retrievalabstractWe describe a new system for searching video databases using free-hand sketched queries. Our query sketches depict both object appearance and motion, and are annotated with keywords that indicate the semantic category of each object. We parse space-time volumes from video to form graph representation, which we match to sketches under a Markov Random Field (MRF) optimization. The MRF energy function is used to rank videos for relevance and contains unary, pairwise and higher-order potentials that reflect the colour, shape, motion and type of sketched objects. We evaluate performance over a dataset of 500 sports footage clips. Rui Hu 0007, Stuart James, Tinghuai Wang, John P. Collomosse |
ICMR | 3 |
| 2013 | State of the "Art": A Taxonomy of Artistic Stylization Techniques for Images and VideoabstractThis paper surveys the field of nonphotorealistic rendering (NPR), focusing on techniques for transforming 2D input (images and video) into artistically stylized renderings. We first present a taxonomy of the 2D NPR algorithms developed over the past two decades, structured according to the design characteristics and behavior of each technique. We then describe a chronology of development from the semiautomatic paint systems of the early nineties, through to the automated painterly rendering systems of the late nineties driven by image gradient analysis. Two complementary trends in the NPR literature are then addressed, with reference to our taxonomy. First, the fusion of higher level computer vision and NPR, illustrating the trends toward scene analysis to drive artistic abstraction and diversity of style. Second, the evolution of local processing approaches toward edge-aware filtering for real-time stylization of images and video. The survey then concludes with a discussion of open challenges for 2D NPR identified in recent NPR symposia, including topics such as user and aesthetic evaluation. Jan Eric Kyprianidis, John P. Collomosse, Tinghuai Wang, Tobias Isenberg 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2012 | Touchcut: Single-touch object segmentation driven by level set methodsabstractIn this paper, we propose an object segmentation algorithm driven by minimal user interactions. Compared to previous user-guided systems, our system can cut out the desired object in a given image with only a single finger touch minimizing user effort. The proposed model harnesses both edge and region based local information in an adaptive manner as well as geometric cues implied by the user-input to achieve fast and robust segmentation in a level set framework. We demonstrate the advantages of our method in terms of computational efficiency and accuracy comparing qualitatively and quantitatively with graph cut based techniques. Tinghuai Wang, John P. Collomosse |
ICASSP | 1 |
| 2012 | Probabilistic Motion Diffusion of Labeling Priors for Coherent Video SegmentationabstractWe present a robust algorithm for temporally coherent video segmentation. Our approach is driven by multi-label graph cut applied to successive frames, fusing information from the current frame with an appearance model and labeling priors propagated forwarded from past frames. We propagate using a novel motion diffusion model, producing a per-pixel motion distribution that mitigates against cumulative estimation errors inherent in systems adopting “hard” decisions on pixel motion at each frame. Further, we encourage spatial coherence by imposing label consistency constraints within image regions (super-pixels) obtained via a bank of unsupervised frame segmentations, such as mean-shift. We demonstrate quantitative improvements in accuracy over state-of-the-art methods on a variety of sequences exhibiting clutter and agile motion, adopting the Berkeley methodology for our comparative evaluation. Tinghuai Wang, John P. Collomosse |
IEEE Trans. Multim. | 1 |
| 2011 | A bag-of-regions approach to sketch-based image retrievalabstractThis paper presents a system for retrieving photographs using free-hand sketched queries. Regions are extracted from each image by gathering nodes of a hierarchical image segmentation into a bag-of-regions (BoR) representation. The BoR represents object shape at multiple scales, encoding shape even in the presence of adjacent clutter. We extract a shape representation from each region, using the Gradient Field HoG (GF-HOG) descriptor which enables direct comparison with the sketched query. The retrieval pipeline yields significant performance improvements over the previous GF-HOG results reliant on single-scale Canny edge maps, and over leading descriptors (SIFT, SSIM) for visual search. In addition, our system enables localization of the sketched object within matching images. Rui Hu 0007, Tinghuai Wang, John P. Collomosse |
ICIP | 2 |
| 2011 | Stylized ambient displays of digital media collections
Tinghuai Wang, John P. Collomosse, Rui Hu 0007, David Slatter, Darryl Greig, Phil Cheatle |
Comput. Graph. | 1 |
| 2010 | Multi-label propagation for coherent video segmentation and artistic stylizationabstractWe present a new algorithm for segmenting video frames into temporally stable colored regions, applying our technique to create artistic stylizations (e.g. cartoons and paintings) from real video sequences. Our approach is based on a multi-label graph cut applied to successive frames, in which the color data term and label priors are incrementally updated and propagated over time. We demonstrate coherent segmentation and stylization over a variety of home videos. Tinghuai Wang, Jean-Yves Guillemaut, John P. Collomosse |
ICIP | 1 |