EDBT 2026 Demo / reviewers in the wild / expert
Gang Wu 0013
dblp:99/6515-13
· DBLP profile ↗
18ranked-venue papers
2as first author
12since 2021 · last 2026
0000-0002-8768-2571ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Atom: Efficient On-Device Video-Language Pipelines Through Modular ReuseabstractRecent advances in video-language models have enabled applications like video retrieval, captioning, and assembly. However, executing such multi-stage pipelines efficiently on mobile devices remains challenging due to redundant model loads and fragmented execution. We introduce Atom, an on-device system that restructures video-language pipelines for fast and efficient execution. Atom decomposes a billion-parameter model into reusable modules, such as the visual encoder and language decoder, and parallely reuses them across subtasks like captioning, reasoning, and indexing. This reuse-centric design eliminates repeated model loading and enables parallel execution, reducing end-to-end latency without sacrificing performance. On commodity smartphones, Atom achieves 27-33% faster execution compared to non-reuse baselines, with only marginal performance drop (≤2.3 Recall@l in retrieval, ≤1.5 CIDEr in captioning). These results position Atom as a practical, scalable approach for efficient video-language understanding on edge devices. Kunjal Panchal, Saayan Mitra, Somdeb Sarkhel, Ishita Dasgupta 0002, Gang Wu 0013, Hui Guan 0001 |
MMSys | 6 |
| 2025 | GUI-Bee: Align GUI Action Grounding to Novel Environments via Autonomous ExplorationabstractGraphical User Interface (GUI) action grounding, mapping language instructions to actionable elements on GUI screens, is important for assisting users in interactive tutorials, task automation, accessibility support, etc. Most recent works of GUI action grounding use large GUI datasets to fine-tune Multimodal Large Language Models (MLLMs). However, the fine-tuning data is inherently limited to specific GUI environments, leading to significant performance degradation in novel environments due to the generalization challenges in the GUI domain. Therefore, we argue that GUI action grounding models should be further aligned with novel environments before deployment to optimize their performance. To address this, we first propose GUI-Bee, an MLLM-based autonomous agent, to collect high-quality, environment-specific data through exploration and then continuously fine-tune GUI grounding models with the collected data. To ensure the GUI action grounding models generalize to various screens within the target novel environment after the continuous fine-tuning, we equip GUI-Bee with a novel Q-value-Incentive In-Context Reinforcement Learning (Q-ICRL) algorithm that optimizes exploration efficiency and exploration data quality. In the experiment, we introduce NovelScreenSpot to test how well the data can help align GUI action grounding models to novel environments. Furthermore, we conduct an ablation study to validate the Q-ICRL method in enhancing the efficiency of GUI-Bee. Handong Zhao, Ruiyi Zhang 0002, Xin Wang 0061, Gang Wu 0013 |
EMNLP | 6 |
| 2025 | Skald: Learning-Based Shot Assembly for Coherent Multi-Shot Video CreationabstractWe present SKALD, a multi-shot video assembly method that constructs coherent video sequences from candidate shots with minimal reliance on text. Central to our approach is the Learned Clip Assembly (LCA) score, a learning-based metric that measures temporal and semantic relationships between shots to quantify narrative coherence. We tackle the exponential complexity of combining multiple shots with an efficient beam-search algorithm guided by the LCA score. To train our model effectively with limited human annotations, we propose two tasks for the LCA encoder: Shot Coherence Learning, which uses contrastive learning to distinguish coherent and incoherent sequences, and Feature Regression, which converts these learned representations into a real-valued coherence score. We develop two variants: a base SKALD model that relies solely on visual coherence and SKALD-text, which integrates auxiliary text information when available. Experiments on the VSPD and our curated MSV3C datasets show that SKALD achieves an improvement of up to 48.6% in IoU and a 43% speedup over the state-of-the-art methods. A user study further validates our approach, with 45% of participants favoring SKALD-assembled videos, compared to 22% preferring text-based assembly methods. Chen-Yi Lu, Md. Mehrab Tanjim, Ishita Dasgupta 0002, Somdeb Sarkhel, Gang Wu 0013, Saayan Mitra, Somali Chaterji |
ICCV | 5 |
| 2023 | Active Context Modeling for Efficient Image and Burst CompressionabstractState-of-the-art compression frameworks usually contain a prediction module and an error context modeling module to reduce redundancy among pixels and improve compression performance. Modern compression algorithms are context adaptive. While adaptive compression algorithms improve over a static context model, they are computationally prohibitive, as the model has to be learned per image during encoding. In this work, we formulate the problem of active context modeling where we train an approximated error context model using an actively selected subset of pixels to significantly speedup the error context modeling while minimizing the impact on compression rate. We investigate the proposed active context modeling framework for both single image compression and burst image compression where the goal is to compress a set of images (usually 6 to 12) captured at a very short time interval between each other. We find that our active context modeling framework is significantly faster than the state-of-the-art while achieving a comparable compression rate. These results indicate the utility of the proposed active context modeling framework for image compression. Gang Wu 0013, Stefano Petrangeli, Ryan Rossi, Viswanathan (Vishy) Swaminathan |
ISM | 2 |
| 2023 | GPU-accelerated Lossless Image Compression with Massive ParallelizationabstractWith the rapid increase of digital content like images or videos nowadays, compression technology contributes more to saving storage or transferring time with large-scale data. While some existing methods already achieved a great compression ratio, they are not applicable to certain live applications under low efficiency. In this work, we use massive parallelization to speed up the SOTA baseline FLIF, including bitwise-equivalent speedup and learning-based speedup. Our method achieves $38.7 \times$ throughputs for encoding and $2.45 \times$ throughputs for decoding, compared to the baseline FLIF. Gang Wu 0013, Viswanathan (Vishy) Swaminathan, Stefano Petrangeli, Tong Yu 0001 |
ISM | 2 |
| 2023 | Content-aware Progressive Image Compression and SyncingabstractProgressive image compression and syncing between devices is an important and challenging problem. When the users are collaboratively editing the same image online, they would expect the changes made by others to be instantly displayed on their side. Since such syncing can be very frequent and usually the image sizes are significantly larger than text data, image live co-editing cannot be easily achieved in the same way as those commonly seen in document co-editing tools. While previous compression techniques like PNG, JPEG and FLIF enable spatially progressive compression, they do not support content-aware compression. Thus, even though the image can be gradually displayed, users cannot prioritize the transmission and display of the most important bits of the image, and often times the resulting pixelation during syncing greatly hurts the user experience. Many existing works on saliency detection can be utilized to provide content awareness. However, many of those techniques are deep-learning-based and it would be computationally prohibitive to directly use them in a latency sensitive scenario like collaborative editing on client devices. In this work, we aim to find a middle ground between a good quality pixel prioritization strategy and extremely fast compression. We start with the pipeline proposed in FLIF and improve it with an entropy-based pixel prioritization strategy, which enables better progressive compression and syncing. Specifically, we modify the traditional Adam interlacing mode [1] to enable an arbitrary pixel transmission order avoiding spatial dependency issues. After constructing the MANIAC tree, we calculate entropy values for each leaf nodes and use them to determine the priority. In addition, we propose to use pixel masks of individual zoom levels to indicate the positions of the transmitted pixels. We further integrate the mask compression algorithm to reduce the communication cost. Through extensive experiments on over 2000 images, we show our proposed method outperforms the baseline methods. Junda Wu, Tong Yu 0001, Gang Wu 0013, Stefano Petrangeli, Handong Zhao, Sungchul Kim, Viswanathan (Vishy) Swaminathan |
ISM | 4 |
| 2022 | One-Pass Algorithms for MAP Inference of Nonsymmetric Determinantal Point ProcessesabstractIn this paper, we initiate the study of one-pass algorithms for solving the maximum-a-posteriori (MAP) inference problem for Non-symmetric Determinantal Point Processes (NDPPs). In particular, we formulate streaming and online versions of the problem and provide one-pass algorithms for solving these problems. In our streaming setting, data points arrive in an arbitrary order and the algorithms are constrained to use a single-pass over the data as well as sub-linear memory, and only need to output a valid solution at the end of the stream. Our online setting has an additional requirement of maintaining a valid solution at any point in time. We design new one-pass algorithms for these problems and show that they perform comparably to (or even better than) the offline greedy algorithm while using substantially lower memory. Aravind Reddy, Ryan Rossi, Zhao Song 0002, Anup B. Rao, Tung Mai, Nedim Lipka, Gang Wu 0013, Eunyee Koh, Nesreen K. Ahmed |
ICML | 7 |
| 2022 | Task-Oriented Near-Lossless Burst CompressionabstractUnlike single images, capturing bursts enables many possible downstream tasks (e.g. superresolution, HDR enhancement) due to the rich information preserved in the consecutive frames. Efficient compression of these bursts is therefore essential given the additional frames to store. In this paper, we propose a novel near-lossless compression method that can preserve the most relevant information in the burst to enable multiple downstream image enhancement tasks, while at the same time reducing the file size. Specifically, we propose a two-bitstream near-lossless compression pipeline that controls the image-space distortion at frame level, and introduce the Lipschitz condition to bound the task-space distortion at burst level. Experiments conducted on a real-world burst dataset confirm the benefit of the proposed solution in terms of rate-distortion both in the burst frame space and the superresolution task space, a popular downstream task in burst processing. Weixin Jiang, Gang Wu 0013, Viswanathan (Vishy) Swaminathan, Stefano Petrangeli, Ryan Rossi, Nedim Lipka |
ISM | 2 |
| 2022 | Show Me What I Like: Detecting User-Specific Video Highlights Using Content-Based Multi-Head AttentionabstractWe propose a method to detect individualized highlights for users on given target videos based on their preferred highlight clips marked on previous videos they have watched. Our method explicitly leverages the contents of both the preferred clips and the target videos using pre-trained features for the objects and the human activities. We design a multi-head attention mechanism to adaptively weigh the preferred clips based on their object- and human-activity-based contents, and fuse them using these weights into a single feature representation for each user. We compute similarities between these per-user feature representations and the per-frame features computed from the desired target videos to estimate the user-specific highlight clips from the target videos. We test our method on a large-scale highlight detection dataset containing the annotated highlights of individual users. Compared to current baselines, we observe an absolute improvement of 2-4% in the mean average precision of the detected highlights. We also perform extensive ablation experiments on the number of preferred highlight clips associated with each user as well as on the object- and human-activity-based feature representations to validate that our method is indeed both content-based and user-specific. Uttaran Bhattacharya, Gang Wu 0013, Stefano Petrangeli, Viswanathan (Vishy) Swaminathan, Dinesh Manocha |
ACM Multimedia | 2 |
| 2021 | From Closing Triangles to Higher-Order Motif Closures for Better Unsupervised Online Link PredictionabstractThis paper introduces higher-order link prediction methods based on the notion of closing higher-order network motifs. The methods are fast and efficient for real-time ranking and link prediction-based applications such as online visitor stitching, web search, and online recommendation. In such applications, real-time performance is critical. The proposed methods do not require any explicit training data, nor do they derive an embedding from the graph data, or perform any explicit learning. Most existing unsupervised methods with the above desired properties are all based on closing triangles (common neighbors, Jaccard similarity, and the ilk). In this work, we develop unsupervised techniques based on the notion of closing higher-order motifs that generalize beyond closing simple triangles. Through extensive experiments, we find that these higher-order motif closures often outperform triangle-based methods, which are commonly used in practice. This result implies that one should consider other motif closures beyond simple triangles. We also find that the best motif closure depends highly on the underlying network and its structural properties. Furthermore, all methods described in this work are fast for link prediction-based applications requiring real-time performance. The experimental results indicate the importance of closing higher-order motifs for unsupervised link prediction. Finally, these new higher-order motif closures can serve as a basis for studying and developing better unsupervised real-time link prediction and ranking methods. Ryan Rossi, Anup B. Rao, Sungchul Kim, Eunyee Koh, Nesreen K. Ahmed, Gang Wu 0013 |
CIKM | 6 |
| 2021 | HighlightMe: Detecting Highlights from Human-Centric VideosabstractWe present a domain- and user-preference-agnostic approach to detect highlightable excerpts from human-centric videos. Our method works on the graph-based representation of multiple observable human-centric modalities in the videos, such as poses and faces. We use an autoencoder network equipped with spatial-temporal graph convolutions to detect human activities and interactions based on these modalities. We train our network to map the activity- and interaction-based latent structural representations of the different modalities to per-frame highlight scores based on the representativeness of the frames. We use these scores to compute which frames to highlight and stitch contiguous frames to produce the excerpts. We train our network on the large-scale AVA-Kinetics action dataset and evaluate it on four benchmark video highlight datasets: DSH, TVSum, PHD2, and SumMe. We observe a 4–12% improvement in the mean average precision of matching the human-annotated highlights over state-of-the-art methods in these datasets, without requiring any user-provided preferences or dataset-specific fine-tuning. Uttaran Bhattacharya, Gang Wu 0013, Stefano Petrangeli, Viswanathan (Vishy) Swaminathan, Dinesh Manocha |
ICCV | 2 |
| 2021 | Provable Distributed Stochastic Gradient Descent with Delayed Updates
Hongchang Gao, Gang Wu 0013, Ryan Rossi |
SDM | 2 |
| 2020 | Structured Policy Iteration for Linear Quadratic RegulatorabstractLinear quadratic regulator (LQR) is one of the most popular frameworks to tackle continuous Markov decision process tasks. With its fundamental theory and tractable optimal policy, LQR has been revisited and analyzed in recent years, in terms of reinforcement learning scenarios such as the model-free or model-based setting. In this paper, we introduce the Structured Policy Iteration (S-PI) for LQR, a method capable of deriving a structured linear policy. Such a structured policy with (block) sparsity or low-rank can have significant advantages over the standard LQR policy: more interpretable, memory-efficient, and well-suited for the distributed setting. In order to derive such a policy, we first cast a regularized LQR problem when the model is known. Then, our Structured Policy Iteration (S-PI) algorithm, which takes a policy evaluation step and a policy improvement step in an iterative manner, can solve this regularized LQR efficiently. We further extend the S-PI algorithm to the model-free setting where a smoothing procedure is adopted to estimate the gradient. In both the known-model and model-free setting, we prove convergence analysis under the proper choice of parameters. Finally, the experiments demonstrate the advantages of S-PI in terms of balancing the LQR performance and level of structure by varying the weight parameter. Youngsuk Park, Ryan Rossi, Gang Wu 0013, Handong Zhao |
ICML | 4 |
| 2019 | Linear Quadratic Regulator for Resource-Efficient Cloud ServicesabstractNo abstract available. Youngsuk Park, Kanak Mahadik, Ryan Rossi, Gang Wu 0013, Handong Zhao |
SoCC | 4 |
| 2019 | Generative Networks for Synthesizing Human Videos in Text-Defined OutfitsabstractGenerating a video from a textual input is a challenging research topic that would have a variety of applications in industries such as retail, e-commerce, online entertainment, education etc. In this paper, we discuss the application of generating videos of a human subject in a desired outfit using an input video of the subject. We present a two stage solution, wherein at the first stage a generative model is learned such that, given the subject's image and a textual description of the outfit, a corresponding image of the subject in the described outfit is synthesized. At the second stage, all the frames of the subject's video are individually processed by the stage 1 model to generate corresponding frames and an optical flow based post processing step is performed to maintain visual coherence across the generated frames. Towards the stage-1 objective, multiple supervised and unsupervised convolutional neural network (CNN) based generative models have been proposed. A novel approach to inject an external masking layer that maintains the structural integrity of the generated images is also presented. We train and test the different methods on the publicly available multi-view clothing image data-set and the performance in videos is showcased on a set of real-world commercial videos. The experiments show the efficacy of our approach in generating images/videos in both low (64 × 64) and high (256 × 256) resolutions. Akshay Malhotra, Viswanathan (Vishy) Swaminathan, Gang Wu 0013, Ioannis D. Schizas |
MMSP | 3 |
| 2019 | Scalable Bid Landscape Forecasting in Real-Time BiddingabstractIn programmatic advertising, ad slots are usually sold using second-price (SP) auctions in real-time. The highest bidding advertiser wins but pays only the second-highest bid (known as the winning price). In SP, for a single item, the dominant strategy of each bidder is to bid the true value from the bidder's perspective. However, in a practical setting, with budget constraints, bidding the true value is a sub-optimal strategy. Hence, to devise an optimal bidding strategy, it is of utmost importance to learn the winning price distribution accurately. Moreover, a demand-side platform (DSP), which bids on behalf of advertisers, observes the winning price if it wins the auction. For losing auctions, DSPs can only treat its bidding price as the lower bound for the unknown winning price. In literature, typically censored regression is used to model such partially observed data. A common assumption in censored regression is that the winning price is drawn from a fixed variance (homoscedastic) uni-modal distribution (most often Gaussian). However, in reality, these assumptions are often violated. We relax these assumptions and propose a heteroscedastic fully parametric censored regression approach, as well as a mixture density censored network. Our approach not only generalizes censored regression but also provides flexibility to model arbitrarily distributed real-world data. Experimental evaluation on the publicly available dataset for winning price estimation demonstrates the effectiveness of our method. Furthermore, we evaluate our algorithm on one of the largest demand-side platforms and significant improvement has been achieved in comparison with the baseline solutions. Aritra Ghosh 0001, Saayan Mitra, Somdeb Sarkhel, Jason Xie, Gang Wu 0013, Viswanathan (Vishy) Swaminathan |
ECML/PKDD (3) | 5 |
| 2017 | Digital content recommendation system using implicit feedback dataabstractMost of existing digital content recommendation systems use explicit feedback data like user's ratings. While such systems rely on user input, those do not sense the context. In this paper, we propose a framework for digital content recommendation using only implicit feedback data (i.e., information collected from session usage without any direct feedback from user), which not only considers interactions among users and contents but also various other implicit information available during a video session. To capture interactions among such attributes, we choose Higher-Order Factorization Machines (HoFM) as our predictor and test our approach on real-world video usage data. In the experiments we explore different possible factors that may affect the performance of HoFM predictor. We observe that increasing the number of sessions of users considered to build the predictor significantly improves prediction accuracy, whereas increasing the order or depth of interactions may not. We also present an application of our work to a video recommendation system. Gang Wu 0013, Viswanathan (Vishy) Swaminathan, Saayan Mitra, Ratnesh Kumar 0001 |
IEEE BigData | 1 |
| 2017 | Context-aware video recommendation based on session progress predictionabstractIn the analysis of digital content consumption, session progress provides a good alternative to using manual ratings for measuring user engagement. A good prediction of session progress is useful for optimizing and personalizing the end-user experience. Most prevalent methods of predicting session progress are based on matrix completion and only consider the interaction among users and videos, while the associated contextual information is usually not used. In this paper, we present our approach for video recommendation, based on session progress prediction and incorporating the context. We test our approach on real-world session progress data, and observe considerable improvement in prediction accuracy achieved by incorporating selected context. Our experiments also show that proper context selection and the number of observed sessions for users are two key factors affecting the prediction accuracy. Gang Wu 0013, Viswanathan (Vishy) Swaminathan, Saayan Mitra, Ratnesh Kumar 0001 |
ICME | 1 |