Back

GPT 5.6 Sol is the best "vision" model OpenAI ever released

11 points2 hoursblog.roboflow.com
kzrdude47 minutes ago

In the third vision bench result, Sol is 100% correct but the expected has 1 error. Seems like an oversight.

In the next bench, Sol looks like it’s correct again but the bboxes are rotated 90 degrees for some reason.

weli56 minutes ago

Anecdotal, opinion:

Gpt is really good in vision stuff, or at least their MoE seems to be really cohesive. From my experience Claude models can be really good at language but the moment they need to look at a picture and decide why the design is not good what parts need improvement it degrades a lot. My easiest benchmark is giving them a screenshot of a feature in my app and tell it "identify non-normative UI blocks and improve readability and consistency". Sol does a great job at re-structuring the page into composable units that build upon each other and the general looks and feels of the app. Claude tends to over-focus one one part while completely forgetting about the rest or the cohesion as a whole.

sscaryterry51 minutes ago

My anecdotal evidence says its still as blind is any other models, it has no taste, no attention to any sort of detail.

howdareme50 minutes ago

How can a vision model have taste?

Razengan50 minutes ago

For 2 weeks I've been trying to get Codex to "outpaint" a wonderful image it generated as placeholder art for a level background.

After I increased the game's resolution, I asked it to increase the image's size while keeping the same scale and existing content, and gosh, it constantly keeps getting something wrong no matter what I tell it, even on Sol Max with the $100 Pro subscription.

An average pixel-artist could have recreated the image and more within 2-3 days.

thatcat47 minutes ago

[delayed]

hn7jmxa7oc58 minutes ago

[dead]