DALL-E 2 Fails to Reliably Capture Common Syntactic Processes






In DALLE2 a picture paints you! More fun poking holes in the current state of natural language understanding in large multimodal models!
I wish the authors included “A man bites a dog.” ;-)
(Previously DALLE2 has toruble with polysemy [1])
Evelina Leivada, Elliot Murphy, and Gary Marcus. 2022. “DALL-E 2 Fails to Reliably Capture Common Syntactic Processes.” arXiv [cs.CL]. arXiv. [2].
Abstract:
Machine intelligence is increasingly being linked to claims about sentience, language processing, and an ability to comprehend and transform natural language into a range of stimuli. We systematically analyze the ability of DALL-E 2 to capture 8 grammatical phenomena pertaining to compositionality that are widely discussed in linguistics and pervasive in human language: binding principles and coreference, passives, word order, coordination, comparatives, negation, ellipsis, and structural ambiguity. Whereas young children routinely master these phenomena, learning systematic mappings between syntax and semantics, DALL-E 2 is unable to reliably infer meanings that are consistent with the syntax. These results challenge recent claims concerning the capacity of such systems to understand of human language. We make available the full set of test materials as a benchmark for future testing.
Originally posted on LinkedIn.
References
[1] Benjamin Han. “Fun with DALL-E 2 and Semantics Leakage.” LinkedIn, 2022. https://www.linkedin.com/posts/benjaminhan_dalle2-semantics-deeplearning-activity-6988955754611314688-RmGD
[2] Evelina Leivada, Elliot Murphy, and Gary Marcus. “DALL-E 2 Fails to Reliably Capture Common Syntactic Processes.” arXiv, 2022. http://arxiv.org/abs/2210.12889