Image-Text-to-Text
Vision-language models that take an image plus a prompt and return text. Powers screenshot understanding, document parsing, chart reading, and GUI automation agents.
Vision-language models that take an image plus a prompt and return text. Powers screenshot understanding, document parsing, chart reading, and GUI automation agents.