HuggingFace task

Image-Text-to-Text

Vision-language models that take an image plus a prompt and return text. Powers screenshot understanding, document parsing, chart reading, and GUI automation agents.

Top models

Agents using Image-Text-to-Text