1 |
Winoground: Probing Vision and Language Models for Visio-Linguistic Compositionality ...
|
|
|
|
Abstract:
We present a novel task and dataset for evaluating the ability of vision and language models to conduct visio-linguistic compositional reasoning, which we call Winoground. Given two images and two captions, the goal is to match them correctly - but crucially, both captions contain a completely identical set of words, only in a different order. The dataset was carefully hand-curated by expert annotators and is labeled with a rich set of fine-grained tags to assist in analyzing model performance. We probe a diverse range of state-of-the-art vision and language models and find that, surprisingly, none of them do much better than chance. Evidently, these models are not as skilled at visio-linguistic compositional reasoning as we might have hoped. We perform an extensive analysis to obtain insights into how future work might try to mitigate these models' shortcomings. We aim for Winoground to serve as a useful evaluation set for advancing the state of the art and driving further progress in the field. The dataset ... : CVPR 2022 ...
|
|
Keyword:
Computation and Language cs.CL; Computer Vision and Pattern Recognition cs.CV; FOS Computer and information sciences
|
|
URL: https://dx.doi.org/10.48550/arxiv.2204.03162 https://arxiv.org/abs/2204.03162
|
|
BASE
|
|
Hide details
|
|
2 |
ANLIzing the Adversarial Natural Language Inference Dataset
|
|
|
|
In: Proceedings of the Society for Computation in Linguistics (2022)
|
|
BASE
|
|
Show details
|
|
3 |
Learning from the Worst: Dynamically Generated Datasets to Improve Online Hate Detection ...
|
|
|
|
BASE
|
|
Show details
|
|
4 |
FLAVA: A Foundational Language And Vision Alignment Model ...
|
|
|
|
BASE
|
|
Show details
|
|
5 |
I like fish, especially dolphins: Addressing Contradictions in Dialogue Modeling ...
|
|
|
|
BASE
|
|
Show details
|
|
6 |
Improving Question Answering Model Robustness with Synthetic Adversarial Data Generation ...
|
|
|
|
BASE
|
|
Show details
|
|
8 |
Gradient-based Adversarial Attacks against Text Transformers ...
|
|
|
|
BASE
|
|
Show details
|
|
10 |
On the Efficacy of Adversarial Data Collection for Question Answering: Results from a Large-Scale Randomized Study ...
|
|
|
|
BASE
|
|
Show details
|
|
11 |
Masked Language Modeling and the Distributional Hypothesis: Order Word Matters Pre-training for Little ...
|
|
|
|
BASE
|
|
Show details
|
|
12 |
Deep Artificial Neural Networks Reveal a Distributed Cortical Network Encoding Propositional Sentence-Level Meaning
|
|
|
|
In: J Neurosci (2021)
|
|
BASE
|
|
Show details
|
|
13 |
Emergent Linguistic Phenomena in Multi-Agent Communication Games ...
|
|
|
|
BASE
|
|
Show details
|
|
14 |
Inferring concept hierarchies from text corpora via hyperbolic embeddings ...
|
|
|
|
BASE
|
|
Show details
|
|
15 |
Inferring concept hierarchies from text corpora via hyperbolic embeddings
|
|
|
|
In: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL 2019) (2019)
|
|
BASE
|
|
Show details
|
|
18 |
Visually Grounded and Textual Semantic Models Differentially Decode Brain Activity Associated with Concrete and Abstract Nouns ...
|
|
|
|
BASE
|
|
Show details
|
|
19 |
Virtual Embodiment: A Scalable Long-Term Strategy for Artificial Intelligence Research ...
|
|
|
|
BASE
|
|
Show details
|
|
20 |
HyperLex: A Large-Scale Evaluation of Graded Lexical Entailment ...
|
|
|
|
BASE
|
|
Show details
|
|
|
|