Last week, OpenAI announced the GPT-5.6 lineup, introducing the Sol, Terra, and Luna models. During the release stream, the squad focused heavy connected computer use, showing models tin of navigating and operating desktop applications. OpenAI highlighted UI agents and elaborate 3D visualizations, but some dangle connected stronger ocular understanding.
To measurement their imagination capabilities, we ran the models done our upcoming VLM benchmark, which we scheme to merchandise successful the adjacent fewer weeks. The benchmark covers communal imagination tasks, including detection, counting, OCR, and information extraction. In this post, we return a person look astatine really GPT-5.6 performs crossed each of them.

Sol is intelligibly the champion imagination exemplary OpenAI has released truthful far. The jump is particularly visible successful entity discovery and counting, wherever GPT-5.5 was acold down the strongest VLMs. Terra and Luna are not arsenic beardown arsenic Sol, but some show meaningful advancement complete GPT-5.5.
Test Sol, Terra, and Luna successful Roboflow Playground and comparison their results pinch models specified arsenic Claude Fable 5 and Gemini 3.5 Flash crossed the aforesaid imagination tasks.
Roboflow Playground
Object Detection
Detection is wherever GPT-5.6 shows the clearest jump. GPT-5.5 scored 13.8 mAP@50 successful our benchmark, while Sol reached 46.2. Terra and Luna followed intimately astatine 44.7 and 43.3, moving entity discovery from a awesome weakness to a applicable capability.

Document layout discovery is 1 of the clearest strengths of GPT-5.6. Sol handled titles, paragraphs, tables, images, and signatures well. Many archive workflows commencement pinch locating the applicable parts of a page earlier OCR aliases information extraction begins.

GPT-5.6 besides performed good connected dense scenes. The pills and eggs examples incorporate galore akin objects packed intimately together, a communal weakness for VLM-based detection. Unlike accepted detectors, VLMs make each people explanation and group of coordinates arsenic text. As entity count grows, the consequence becomes longer and the consequence of missed objects, duplicates, aliases coordinate errors increases. Despite this, Sol detected astir objects crossed some scenes.

For the champion discovery results, punctual GPT-5.6 models to return absolute XYXY coordinates successful image pixels. This differs from Gemini 3.5 Flash, which performed champion pinch YXYX coordinates normalized to a 0–1000 range. Using the incorrect coordinate format reduced GPT-5.6 discovery capacity by astir 15 mAP points successful our benchmark.
In a fewer cases, GPT-5.6 Sol returned boxes successful seemingly random parts of the image. Many had nary overlap, aliases almost nary overlap, pinch the crushed truth. Instead of matching the visible objects, the boxes often formed unnatural layouts, specified arsenic consecutive rows aliases evenly spaced groups.


We shared those examples pinch OpenAI. Their squad confirmed that Sol becomes little unchangeable connected images astir 2,000 by 2,000 pixels aliases larger, particularly astatine little reasoning effort. Higher reasoning effort improves stability, but besides increases token use, latency, and cost. Resizing aliases cropping ample images earlier sending them to the OpenAI API is the astir applicable workaround.
Object Counting
Counting improved crossed the afloat GPT-5.6 lineup. Sol scored 73.0% successful our benchmark, up from 64.9% for GPT-5.5, while Terra and Luna reached 67.6% and 66.2%. Luna, the cheapest exemplary successful the lineup, still outperformed the erstwhile OpenAI baseline.

As portion of the benchmark, we tested cases requiring much than spotting objects and returning a total. Sol counted heavy overlapping metallic brackets, a difficult lawsuit for some accepted entity detectors and VLMs. Sol besides counted slug holes only wrong selected scoring zones, showing an knowing of some which objects to count and wherever the norm applied.


Blister packs proved overmuch harder. In abstracted prompts, we asked Sol to count the quiet slots and the pills still sealed wrong the package. The repeated layout, reflections, and mini ocular differences betwixt filled and quiet slots made some tasks difficult.


The abnormal candy illustration exposed a different type of failure. Sol gave the incorrect count, though it is unclear whether the exemplary miscounted the candies aliases misunderstood the target category.

OCR and Data Extraction
OCR capacity stayed adjacent to GPT-5.5. Sol achieved a 90.7% mean similarity score, only 0.5 points down GPT-5.5 astatine 91.2%, while Terra and Luna reached 88.8% and 88.4%. The spread was larger successful matter extraction, wherever Sol scored 82.5% compared pinch 87.6% for GPT-5.5. Luna and Terra followed astatine 81.4% and 79.4%.


As portion of the benchmark, we separated afloat transcription from targeted extraction. OCR asks the exemplary to transcribe each visible text, while matter extraction asks for a circumstantial portion of information. Sol performed good connected handwritten notes successful some settings, producing a afloat transcription successful 1 lawsuit and extracting a requested day successful another.


Sol performed good connected matter embedded successful analyzable ocular scenes. It publication a tyre size series printed on the curved aboveground of a dirty, worn tire. In different example, it extracted the unrecorded people from a lucky broadcast and returned the reply successful the requested format, testing some ocular reference and instruction following.


Some simple-looking extraction tasks still failed. Sol could not publication the expiration day printed connected a blister pack. The matter was small, vertical, debased contrast, and affected by reflections, which whitethorn explicate the error.

Trade-offs
The imagination gains travel pinch higher token usage crossed the GPT-5.6 lineup. The quality matters little successful mini tests, but becomes much important astatine scale, wherever token measurement straight increases processing costs.

Sol averaged adjacent to 10 seconds per image successful our benchmark. Terra reduced that to astir 6 seconds, while Luna vanished successful somewhat complete 5 seconds. Luna offers the strongest latency-quality equilibrium successful the lineup, pinch velocity adjacent to Gemini 3.5 Flash while still outperforming GPT-5.5 connected discovery and counting.

In our benchmark, Sol costs astir 2.5 cents per image, making it the 2nd astir costly exemplary aft Claude Fable 5. Terra reduced the mean costs to astir 1 cent per image, while Luna costs little than 0.5 cents.

At 0.8 cents per image, Gemini 3.5 Flash is overmuch cheaper than Sol while still starring our discovery and counting benchmarks. This makes it a beardown action for data-intensive workloads wherever costs scales crossed ample image batches. Roboflow Playground lets you trial Sol, Terra, and Luna alongside Claude Fable 5, Gemini 3.5 Flash, and different VLMs connected the aforesaid tasks.
Takeaways
With GPT-5.6, OpenAI is overmuch person to the starring VLMs than before. Detection moved from a anemic constituent to a usable capability, and counting improved crossed the afloat exemplary family.
There are still clear limits. Gemini 3.5 Flash remains a amended applicable prime for high-volume discovery and counting successful our benchmark, particularly astatine its price.
GPT-5.6 shows OpenAI is now taking imagination overmuch much seriously. Sol still has flaws, particularly astir cost, latency, and immoderate unstable discovery cases, but the advancement is difficult to ignore. For agents, surface understanding, archive workflows, and ocular reasoning, this merchandise makes OpenAI a overmuch stronger action than before.
English (US) ·
Indonesian (ID) ·