A quote from Alex Albert (Anthropic)

Yeah, unfortunately vision prompting has been a tough nut to crack. We've found it's very challenging to improve Claude's actual "vision" through just text prompts, but we can of course improve its reasoning and thought process once it extracts info from an image.

In general, I think vision is still in its early days, although 3.5 Sonnet is noticeably better than older models.

— Alex Albert (Anthropic)

Posted 10th July 2024 at 6:56 pm

Simon Willison’s Weblog

Recent articles