xAI's Imagine Image 2.0 Debuts Behind GPT Image 2: No API, No Pipeline Yet
On August 7, 2026 the fast variant of xAI's new image model debuted second on both Arena boards, behind OpenAI's GPT Image 2 on each. Here is why that does not change what an ad team should be generating creative with this week, and the three things that would.
Mauricio Valdivia
·11 min

Second on the board, and nowhere in your pipeline
A media buyer drops a leaderboard screenshot into the team channel with one line under it: should we switch? The screenshot is real. The board is real. The honest answer takes about ninety seconds, because the model sitting in second place is one nobody on that team can call.
On August 7, 2026, xAI shipped Imagine Image 2.0 into Grok's web and mobile apps, and its fast variant debuted second on both of Arena's image boards. On Image Edit it scored 1439 Elo, behind OpenAI's gpt-image-2 (medium) at 1463. On Text-to-Image it scored 1320, behind gpt-image-2 at 1380. Two silver medals in one day.
Then the part that decides whether any of it reaches your ad account. There is no API. xAI says access is coming soon and attaches no date, and its own developer price list still sells exactly two image models, neither of which is this one. Meanwhile gpt-image-2 is first on both boards, and GPT Image 2 is the model the Novoads image-to-ad and product-to-ad flows already run. So the short answer to "does this change what I generate ad creative with" is no, not yet. The long answer is the useful one, because it tells you what would.
What actually landed on August 7
Three details in this launch are load-bearing, and all three get flattened in the aggregate coverage. They are worth ten minutes now because they will keep mattering for the next model launch too.
The entry that ranks is not the model in the app
xAI launched Imagine Image 2.0 as a new Quality Mode on grok.com/imagine and in the Grok iOS and Android apps. The row sitting at number two on Arena is labelled grok-imagine-image-2.0 (low), the fast tier of that family. The Quality Mode consumers actually received is not the row that ranked.
OpenAI's leading row is gpt-image-2 (medium), which is also a tier rather than a ceiling. So the comparison being celebrated is one speed tier of a new model against one middle tier of an incumbent, which is a legitimate result and a narrower one than "second best in the world." Any headline that says Imagine Image 2.0 ranks second, with no qualifier, has quietly promoted a variant into a product.
Preliminary is not a formality
Both Grok rows carry Arena's Preliminary flag, and the vote counts explain why.
| Board, as of Aug 7, 2026 | First | Second | Votes, first vs second |
|---|---|---|---|
| Image Edit | gpt-image-2 (medium), 1463 | grok-imagine-image-2.0 (low), 1439 | 184,189 vs 5,931 |
| Text-to-Image | gpt-image-2 (medium), 1380 | grok-imagine-image-2.0 (low), 1320 | 69,194 vs 2,722 |
A model with 5,931 votes and a model with 184,189 votes are not being measured with the same instrument. The confidence intervals say so out loud: plus or minus 8 on the challenger, plus or minus 4 on the leader in Image Edit, and plus or minus 12 against plus or minus 5 in Text-to-Image. A day-one placement is a first impression, not a verdict, and xAI itself date-stamps the claim rather than stating it flatly.
On one board, second is not separated from third
Text-to-Image is the thinner of the two results. The second row reads 1320 with a plus or minus 12 interval, and the third row, Reve 2.1, reads 1302 with plus or minus 8. Those ranges overlap. Statistically, second and third on that board are not yet distinguishable, which means the ordering can change without the model changing at all.
None of this makes the release unimpressive. It makes the number provisional, and provisional numbers are exactly the ones that end up quoted for months after they stopped being true.

The part most coverage will skip: there is no API
Leaderboard stories are easy to write and easy to read. Availability stories are neither, which is why the most decision-relevant fact in this launch is the one that travels least.
What xAI's price list actually sells today
xAI's own developer pricing page lists the Imagine API image models it sells. There are two. grok-imagine-image-quality runs $0.05 per image at 1K and $0.07 at 2K. grok-imagine-image runs $0.02 per image at both 1K and 2K. There is no Imagine Image 2.0 row.
That absence is meaningful rather than merely quiet. This is the complete page where a purchasable image model would appear, which is the only condition under which a missing entry tells you anything. It is not an omission somewhere in a docs tree. It is the price list.
The models you can call rank 13th and 16th
Here is the sentence worth keeping from this entire launch cycle. On the same Image Edit board where the unreleased variant placed second, the two xAI image models a developer can actually buy sit at 13 and 16: grok-imagine-image-quality at 1362, and grok-imagine-image at 1330. On Text-to-Image, grok-imagine-image-quality sits at 14 with 1228 Elo on 48,786 votes.
So the headline number and the purchase options are two different populations. The gap between them, roughly a hundred Elo points on Text-to-Image, is precisely the value that has not shipped to developers yet. If your workflow lives in an API, the honest xAI number today is 14th place, not second.
Coming soon has no date on it
xAI's stated position is that API access for Imagine Image 2.0 is coming soon. That is a real commitment, and it is unusable for planning. Four things a creative pipeline needs are all still missing:
- A model ID to write into a config, next to the two that already exist.
- A published developer price, so a month of variations is a number rather than a guess.
- A rate limit, so a batch can be designed instead of discovered.
- A date, so a migration can be scheduled against something.
Which is the difference between a model that exists and a model you can build on. An ad team does not consume an image model by admiring it. It consumes one by calling it forty times on a Tuesday, in the ratio the media plan asked for, against a product photo that has to stay recognisable. Everything in that sentence needs an endpoint.
An Elo is not an ad-creative score
Suppose the API lands next month with sane pricing. The leaderboard still would not have answered the question an ad team is actually asking, and this is worth separating from the availability problem because it survives it.
What the vote actually measures
An Arena score aggregates blind pairwise human preference on prompts the voters chose to write. It is a genuinely good instrument for general taste, and it is the reason the boards are worth reading at all. It is also, structurally, a popularity measure over a prompt distribution that has nothing to do with your category.
What it does not measure is most of what an ad reviewer checks:
- whether the serum label stayed legible and said what the real label says
- whether the packaging shade survived the render
- whether the headline reads at thumb speed on a phone in daylight
- whether the fortieth asset in a campaign still matches the first
Those are the things that decide whether a creative ships. We went through the same distinction when Meta pitched Muse Image at advertisers, and again when Microsoft's MAI Image 2.5 Pro arrived with hero-imagery language and a one-megapixel ceiling: the announcement and the constraint that stops you on a Thursday are usually written on different pages.
The typography claim is the one to test yourself
xAI positions the model on instruction-following accuracy, clean typography and layout in complex visuals, and consistency across multiple generations. Those are exactly the right things to claim for ad work, and exactly the claims a leaderboard Elo cannot confirm.
Text rendering is where image models diverge most sharply and where marketing language is least reliable. Ideogram 4 built its whole release around in-image text, with explicit bounding-box layout and palette control, and Seedream 5.0 Pro leans on multilingual text rendering for teams shipping one offer into several markets. Those are structural commitments you can inspect. A preference score is not.
Pin every mark, or the model writes its own
The failure mode we hit most often has nothing to do with resolution. An image model handed a product without explicit instruction will invent brand copy, and the inventions are always plausible:
- a claim on the label that reads like marketing but that nobody approved
- a tagline the brand has never used
- a certification badge that does not exist
- a variant name lifted from the general shape of the category
It all looks correct at a glance, which is what makes it dangerous, and it survives right up until legal reads the ad.
The fix is not a better model. It is pinning every mark in the prompt, then transcribing what came back from your own render rather than from the gallery. That discipline is the same one that keeps an AI actor consistent across ad variations, and it is why the prompt craft around a model tends to matter more than its rank, right up until the model absorbs the workaround and turns it into a feature. A new number two on a public board does not, on its own, change how anyone works.

The model at the top of both boards is already drawing your ads
The quiet result in this story is not the challenger. It is that gpt-image-2 (medium) held first place on both boards on the same day, on the largest vote counts on either page.
What GPT Image 2 costs inside Novoads
That model is not a thing to go and evaluate. It is what the Novoads image-to-ad and product-to-ad flows run, at fal medium quality, at a flat 0.3 credits per image. One number, one balance, no separate provider key, and no per-token rate to reverse-engineer before you can answer what a batch of variations costs. Nano Banana Pro, Seedream 5 Pro, Seedream 5 Lite and Reve 2.1 sit in the same catalog when a job wants a different engine.
The economics matter more than the ranking here. Creative testing is a volume activity, and testing creative properly means generating far more candidates than you will ever run. A flat per-image price makes that a multiplication you can do in your head.
Do the multiplication once and the difference gets concrete. A week of image testing at four concepts, five variations each, in two placement ratios is 40 renders. At 0.3 credits an image that is 12 credits, which is a quarter of what the entry plan grants in a month, and you knew the number before you started. Now price the same week against an endpoint that bills per output token, and the same question turns into a forecast you cannot settle until the invoice arrives. That is not a rounding difference between vendors. It is the difference between a budget and a guess.
Why the ratio list matters more than the Elo
The image-to-ad flow exposes six aspect ratios: 4:5, 9:16, 1:1, 16:9, 2:3 and 21:9. That list is a more practical spec than any Elo on either board, because a creative in the wrong shape is not a slightly worse creative, it is one the placement will not take. Check your ratio requirements per platform before you generate rather than after, whichever model you end up using.
What would actually change the answer
None of the above is a verdict on the model. It is a verdict on the timing. Three specific things would move it, and it is worth writing them down now so the decision is not made on vibes when they arrive.
The three signals to watch
- A model ID and a published price on xAI's own pricing page. Not an announcement, not a waitlist. A row you can put in a config next to the two that are already there.
- The vote counts crossing into the tens of thousands, with the Preliminary flag gone. At that point the placement is a measurement rather than a first impression, and it may well be higher or lower than 1439.
- A benchmark on your own product, not the prompt distribution. Ten of your real SKUs, your real labels, your real headline, in the ratio you actually run. That is the only test that answers your question.
What to do this week instead
Open Grok Imagine and look at the output if you are curious. It costs you ten minutes and it is a genuinely strong model. Then go back to the pipeline you can automate, because a creative operation is a throughput problem, and throughput needs an endpoint.
It is worth noticing the contrast with the other kind of launch news. Ad-platform changes tend to arrive with a switch-over date attached, which is what makes them plannable: our note on the ChatGPT Ads matching and carousel changes is the shape of a story you can put in a calendar. A model with no endpoint is a story you can only react to.
If you want a useful thing to change this week, change the input rather than the model. Most weak AI creative is not weak because the renderer lost a benchmark. It is weak because nobody wrote an angle, which is the quality gap that gets blamed on models roughly every time a new one launches.
Which document answers which question
Read the leaderboard when you want to know:
- which labs are currently competitive at all
- whether a family improved between versions
- which models are worth an afternoon of hands-on testing
Read the vendor's price list when you want to know:
- what you can call tomorrow morning
- what a month of variations will cost
- how many requests a batch is allowed to make
- whether a migration can be scheduled at all
Only one of those two documents can be turned into a media plan, and it is not the one that made the headlines this week.

How Novoads solves the model-churn problem
There is a new frontier image model roughly every two weeks now. Chasing each one has a real cost, and almost none of it is judgment: another account, another key, another billing unit, another integration to maintain when the model gets deprecated.
One balance, and someone else's integration problem
In Novoads you upload a product image, and the image-to-ad and product-to-ad flows generate the creative on GPT Image 2 at medium quality for 0.3 credits an image, in whichever of the six ratios the placement wants. When a model earns a slot, it appears in the picker on the same balance. When it does not, nothing on your side changes. Imagine Image 2.0 is not in that picker today, and saying so plainly is better than implying an integration that cannot exist without an endpoint.
Where the still goes next
Then the step a still cannot take on its own. Write or auto-generate a script, pick an AI actor whose age, gender and accent match the audience, and the approved image becomes the opening frame of a vertical clip with synthetic voice, lip-sync and captions. Our guide to making product videos with AI walks that loop end to end, and the same handoff drives the route from one key frame to a finished animated spot.
Trying it costs $1 for three days of access, which then becomes the $49-a-month Inicial plan. That first charge grants 10 credits, enough for about one video. Cancel anytime.
Ship on what you can call
Imagine Image 2.0 is a serious release from a serious lab, and the low variant landing second on two boards on day one is a real result that deserves the attention it got. It also has no model ID, no developer price, no rate limit and no date, which means for anyone running ads it is a preview of a decision rather than a decision.
The pattern is going to repeat, because the leaderboard is the part of a launch that is easy to publish and availability is the part that is easy to skip. So keep the two questions apart. A benchmark tells you what a model can do in someone else's prompt. A price list tells you what you can build on Tuesday. When those two disagree, the price list is the one that decides what ships.
Frequently Asked Questions
What is xAI Imagine Image 2.0?
Imagine Image 2.0 is xAI's image generation and editing model, launched as a new Quality Mode on grok.com/imagine and in Grok's iOS and Android apps. xAI positions it on instruction-following accuracy, clean typography and layout in complex visuals, and consistency across multiple generations. It is a consumer product today rather than a developer one.
Is Imagine Image 2.0 really the second best image model in the world?
It is more precise than that. On Arena's boards dated August 7, 2026, the entry in second place is grok-imagine-image-2.0 (low), the fast variant, at 1439 Elo on Image Edit and 1320 on Text-to-Image, behind gpt-image-2 (medium) at 1463 and 1380. Both Grok rows are flagged Preliminary on a few thousand votes against tens of thousands for the leader, so the placement is provisional and can move.
Can I use Imagine Image 2.0 through an API?
Not as of August 8, 2026. xAI says API access is coming soon and gives no date. Its own developer pricing page lists two Imagine API image models, grok-imagine-image and grok-imagine-image-quality, with no 2.0 entry, so there is no model ID to call and no published developer price to budget against.
How much do xAI's callable image models cost?
xAI's developer pricing page lists grok-imagine-image at $0.02 per image at both 1K and 2K, and grok-imagine-image-quality at $0.05 per image at 1K and $0.07 at 2K. Those are the two image models a developer can actually reach today. Neither is the model that placed second on the leaderboards.
Does a leaderboard Elo tell me which model makes better ads?
No. An Arena score aggregates blind human preference between two outputs on prompts the voters chose, which is a general taste signal rather than an ad-performance one. It does not measure whether your product label survived, whether the headline is legible at thumb speed, or whether the fortieth asset in a campaign still matches the first. Those you test on your own product shots.
Which image model does Novoads use for ad creative?
The image-to-ad and product-to-ad flows run GPT Image 2 at fal medium quality, at a flat 0.3 credits per image, across six aspect ratios: 4:5, 9:16, 1:1, 16:9, 2:3 and 21:9. Nano Banana Pro, Seedream 5 Pro, Seedream 5 Lite and Reve 2.1 are also in the catalog. Imagine Image 2.0 is not, because there is nothing to integrate yet.
Key Takeaways
- The entry ranked second on both Arena boards is grok-imagine-image-2.0 (low), the fast variant, not the Quality Mode that xAI actually shipped into the Grok apps. Any headline that drops that qualifier is overstating the result.
- Both placements are provisional. Arena flags the Grok rows Preliminary at 5,931 votes on Image Edit and 2,722 on Text-to-Image, against gpt-image-2's 184,189 and 69,194, and on Text-to-Image the second and third intervals overlap.
- There is no API. xAI says access is coming soon with no date, and its own developer price list still sells only grok-imagine-image at $0.02 per image and grok-imagine-image-quality at $0.05 at 1K and $0.07 at 2K.
- The xAI image models you can actually call rank 13th and 16th on Image Edit and 14th on Text-to-Image. The number that made the headlines and the products a developer can buy are not the same thing.
- GPT Image 2 is number one on both boards as of August 7, 2026, and it is the model Novoads already runs for image-to-ad and product-to-ad at fal medium quality, at 0.3 credits per image across six placement ratios.




