VLDB 2026 Research / reviewers in the wild / expert
Qing Sun 0001
dblp:34/7845-1
· DBLP profile ↗
4ranked-venue papers
3as first author
0since 2021 · last 2018
0009-0005-9647-882XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Language models and text generation · 29% Vision and language · 23% Deep learning architectures and training · 13% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% |
Topics — the 11 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language
image captioning |
0.4 | 2 | 2018 | Bidirectional Beam Search: Forward-Backward Inference in Neural Sequence Models for Fill-in-the-Blank Image Captioning · CVPR 2017 Diverse Beam Search for Improved Description of Complex Scenes · AAAI 2018 |
Natural language and speech › Language models and text generation
decoding |
0.3 | 1 | 2018 | Diverse Beam Search for Improved Description of Complex Scenes · AAAI 2018 |
Natural language and speech › Language models and text generation › decoding
bidirectional decoding |
0.3 | 1 | 2017 | Bidirectional Beam Search: Forward-Backward Inference in Neural Sequence Models for Fill-in-the-Blank Image Captioning · CVPR 2017 |
Machine learning › Deep learning architectures and training › sequence modeling
neural sequence models |
0.3 | 1 | 2017 | Bidirectional Beam Search: Forward-Backward Inference in Neural Sequence Models for Fill-in-the-Blank Image Captioning · CVPR 2017 |
Machine learning › Efficient and distributed learning
active learning |
0.2 | 1 | 2015 | Active learning for structured probabilistic models with histogram approximation · CVPR 2015 |
Machine learning › Probabilistic and Bayesian machine learning › structured prediction
conditional random field |
0.2 | 1 | 2015 | Active learning for structured probabilistic models with histogram approximation · CVPR 2015 |
Computer vision › Image recognition and object detection › object detection
object proposal generation |
0.2 | 1 | 2015 | SubmodBoxes: Near-Optimal Search for a Set of Diverse Object Proposals · NIPS 2015 |
Mathematical optimization › submodular optimization › submodular maximization
monotone submodular maximization |
0.2 | 1 | 2015 | SubmodBoxes: Near-Optimal Search for a Set of Diverse Object Proposals · NIPS 2015 |
Mathematical optimization › submodular optimization
submodular maximization |
0.2 | 1 | 2015 | SubmodBoxes: Near-Optimal Search for a Set of Diverse Object Proposals · NIPS 2015 |
Natural language and speech › Machine translation
neural machine translation |
0.1 | 1 | 2018 | Diverse Beam Search for Improved Description of Complex Scenes · AAAI 2018 |
Computer vision › Vision and language › vision-language generation
visual question generation |
0.1 | 1 | 2018 | Diverse Beam Search for Improved Description of Complex Scenes · AAAI 2018 |
Methods — techniques the papers use, named apart from their topics
beam search · 0.6non-maximal suppression · 0.4lazy greedy · 0.4branch-and-bound · 0.4recurrent neural network · 0.3diverse beam search · 0.3LSTM · 0.3bidirectional beam search · 0.3histogram approximation · 0.2entropy approximation · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2018 | Diverse Beam Search for Improved Description of Complex ScenesabstractA single image captures the appearance and position of multiple entities in a scene as well as their complex interactions. As a consequence, natural language grounded in visual contexts tends to be diverse---with utterances differing as focus shifts to specific objects, interactions, or levels of detail. Recently, neural sequence models such as RNNs and LSTMs have been employed to produce visually-grounded language. Beam Search, the standard work-horse for decoding sequences from these models, is an approximate inference algorithm that decodes the top-B sequences in a greedy left-to-right fashion. In practice, the resulting sequences are often minor rewordings of a common utterance, failing to capture the multimodal nature of source images. To address this shortcoming, we propose Diverse Beam Search (DBS), a diversity promoting alternative to BS for approximate inference. DBS produces sequences that are significantly different from each other by incorporating diversity constraints within groups of candidate sequences during decoding; moreover, it achieves this with minimal computational or memory overhead. We demonstrate that our method improves both diversity and quality of decoded sequences over existing techniques on two visually-grounded language generation tasks---image captioning and visual question generation---particularly on complex scenes containing diverse visual content. We also show similar improvements at language-only machine translation tasks, highlighting the generality of our approach. Ashwin K. Vijayakumar, Michael Cogswell, Ramprasaath R. Selvaraju, Qing Sun 0001, Stefan Lee, David Crandall, Dhruv Batra |
AAAI | 4 |
| 2017 | Bidirectional Beam Search: Forward-Backward Inference in Neural Sequence Models for Fill-in-the-Blank Image CaptioningabstractWe develop the first approximate inference algorithm for 1-Best (and M-Best) decoding in bidirectional neural sequence models by extending Beam Search (BS) to reason about both forward and backward time dependencies. Beam Search (BS) is a widely used approximate inference algorithm for decoding sequences from unidirectional neural sequence models. Interestingly, approximate inference in bidirectional models remains an open problem, despite their significant advantage in modeling information from both the past and future. To enable the use of bidirectional models, we present Bidirectional Beam Search (BiBS), an efficient algorithm for approximate bidirectional inference. To evaluate our method and as an interesting problem in its own right, we introduce a novel Fill-in-the-Blank Image Captioning task which requires reasoning about both past and future sentence structure to reconstruct sensible image descriptions. We use this task as well as the Visual Madlibs dataset to demonstrate the effectiveness of our approach, consistently outperforming all baseline methods. Qing Sun 0001, Stefan Lee, Dhruv Batra |
CVPR | 1 |
| 2015 | Active learning for structured probabilistic models with histogram approximationabstractThis paper studies active learning in structured probabilistic models such as Conditional Random Fields (CRFs). This is a challenging problem because unlike unstructured prediction problems such as binary or multi-class classification, structured prediction problems involve a distribution with an exponentially-large support, for instance, over the space of all possible segmentations of an image. Thus, the entropy of such models is typically intractable to compute. We propose a crude yet surprisingly effective histogram approximation to the Gibbs distribution, which replaces the exponentially-large support with a coarsened distribution that may be viewed as a histogram over M bins. We show that our approach outperforms a number of baselines and results in a 90%-reduction in the number of annotations needed to achieve nearly the same accuracy as learning from the entire dataset. Qing Sun 0001, Ankit Laddha, Dhruv Batra |
CVPR | 1 |
| 2015 | SubmodBoxes: Near-Optimal Search for a Set of Diverse Object ProposalsabstractThis paper formulates the search for a set of bounding boxes (as needed in object proposal generation) as a monotone submodular maximization problem over the space of all possible bounding boxes in an image. Since the number of possible bounding boxes in an image is very large $O(#pixels^2)$, even a single linear scan to perform the greedy augmentation for submodular maximization is intractable. Thus, we formulate the greedy augmentation step as a Branch-and-Bound scheme. In order to speed up repeated application of B\&B, we propose a novel generalization of Minoux’s ‘lazy greedy’ algorithm to the B\&B tree. Theoretically, our proposed formulation provides a new understanding to the problem, and contains classic heuristic approaches such as Sliding Window+Non-Maximal Suppression (NMS) and and Efficient Subwindow Search (ESS) as special cases. Empirically, we show that our approach leads to a state-of-art performance on object proposal generation via a novel diversity measure. Qing Sun 0001, Dhruv Batra |
NIPS | 1 |