Mirage Avatar X costs 71% more per second on fal than from the vendor
Mirage lists Avatar X at $0.175 per second on its own API while fal lists the same model at $0.3 per second, and the billing increment underneath those numbers changes the answer again.
Mauricio Valdivia
·11 min
One model, two price tags, one 71% higher
It is late, the sprint board says "ship the avatar endpoint," and there is a code sample on the screen with a copy button next to it. You paste the key, run the sample, watch a talking head come back, and move on. Three weeks later somebody in finance asks why the video line is what it is.
That is the whole story here, and it has a number attached. Mirage's own API pricing page lists its Avatar X model at $0.175 per second. fal, which resells the same model as a partner endpoint, tells you on the model page that your request will cost $0.3 per second. Same model, same vendor behind it, both prices live and public on August 11, 2026, and one of them is about 71% higher per second than the other. Underneath that there is a second, quieter difference in how each side counts the seconds, and on some clip lengths it flips the answer entirely.
The company behind the model is one you probably already know
Before the prices, the entity, because this one trips people up in search and in procurement.
NOCAP, Inc., which used to trade as Captions
Mirage and Captions are the same company. The footer on mirage.app reads "© 2026 NOCAP, Inc. d/b/a Mirage. All rights reserved.", and the same footer sits on captions.ai. TechCrunch covered the rename in September 2025. The Captions app, the consumer video tool a lot of marketers have had on a phone since 2023, is a product of that company; Mirage is the model and research side of it.
This matters practically, not just as trivia. It is why the API documentation and the pricing page for a model branded "Mirage Avatar X" still live at captions.ai, and why a procurement search for "Mirage pricing" can come back looking thin. If you have ever evaluated AI avatar video generators and put Captions in one column and Mirage in another, they belong in the same one.
What Avatar X actually produces
The model takes a script or a driving audio track plus a reference, and returns one lip-synced MP4 with the voice already inside it. There is no separate audio stem to mix. Mirage's docs describe the intent as generating the whole performance in a single pass, keeping identity consistent across a full segment while carrying emotion, laughter, micro-expressions and nonverbal movement, rather than assembling those from separate stages.
fal exposes two arms of it, priced identically:
- Text to video. A script of 50 to 1,500 characters, plus one of ten stock characters from a fixed catalog. Each is offered in a vertical and a horizontal cut, so the aspect ratio is chosen by picking
JasmineorJasmine (16:9)rather than by setting a field. - Reference to video. A driving audio clip plus a video reference. Supply both and they replace the stock avatar entirely.
The second arm is the closer match to how Mirage describes the model on its own pricing page: an image and audio to video model. Note what is missing from both. There is no voice parameter. The voice arrives bundled with the stock character, or it is cloned from the audio reference you hand over, of which fal says only the first 60 seconds are used.
The launch was two weeks before the listing
Avatar X was announced on July 28, 2026, in a post on Mirage's own blog that opens "Today we're announcing Mirage Avatar X, our newest avatar model." The fal listing is the new thing: fal's catalog data stamps both endpoints public with a publish time on August 11, 2026.
So this is not a launch story. It is a distribution story, and distribution stories are where prices diverge.

Two published prices for the same endpoint
Both numbers are quotable, both are current, and neither is a negotiated rate hiding behind a sales call.
What Mirage charges on its own API
Mirage's API pricing page lists one line for the model: expressive human video generation, image plus audio to video, at "$0.175 per second (6-second increments)". The same page states the rule that governs it: you are billed for the length of the generated output video, rounded up to the nearest 6-second increment.
Two facts in one sentence. A rate, and a unit. Most people copy the rate.
Getting that price requires a Mirage platform account and an API key created in the dashboard. There is no enterprise gate on it and no "contact sales" wall, which is what makes it a genuine list price rather than a teaser.
What fal charges for the same model
On fal's model page for the text-to-video arm, under the request panel, sits a single sentence: your request will cost $0.3 per second. The identical line appears on the reference-to-video arm. fal publishes no minimum increment alongside it.
Why "the same model" is not an inference
This is worth pinning down, because reselling stories usually founder on whether the two things are actually the same thing. fal's own catalog data answers it. The entries carry "modelLab":"Mirage" and "modelFamily":"Avatar X", the endpoints are namespaced under mirage-api, and the description fal publishes is Mirage's own marketing sentence reproduced verbatim, down to the claim about identity preservation. fal's avatar documentation links back to Mirage's own avatar catalog on captions.ai for the character names.
It is a partner listing of the vendor's API. Not a fine-tune, not a clone, not a similarly named competitor.
The billing increment is the second price
Here is the part that separates a headline from a purchase decision, and it is the reason a per-second rate on its own is close to meaningless.
A six-second block is not a per-second rate
Mirage bills in 6-second blocks, rounded up. So a 7-second output is billed as 12 seconds. A 13-second output is billed as 18. The effective per-second cost of a Mirage render is $0.175 only when the output lands exactly on a multiple of six, and it is higher, sometimes much higher, everywhere else.
fal's listing states a per-second price and no minimum. Those are two different products wearing the same kind of price tag.
Where the gap is exactly 71%, and where it is not
Run both meters against real clip lengths and the single headline number turns into a curve. Prices below are computed from each surface's own published rate and rule.
| Output length | Mirage billed as | Mirage cost | fal cost | fal premium |
|---|---|---|---|---|
| 6 seconds | 6 s | $1.05 | $1.80 | +71% |
| 6.5 seconds | 12 s | $2.10 | $1.95 | 7% cheaper |
| 12 seconds | 12 s | $2.10 | $3.60 | +71% |
| 15 seconds | 18 s | $3.15 | $4.50 | +43% |
| 60 seconds | 60 s | $10.50 | $18.00 | +71% |
The 71% is real and it is the ceiling. It holds whenever your output is a clean multiple of six, which for a scripted read it very rarely is, because you do not control the exact length of generated speech to the tenth of a second. On a 15-second clip, the length most short-form ad workflows actually target, the gap is closer to 43%.
The narrow window where the reseller is cheaper
There is a band, just past each 6-second boundary, where fal is the better buy. Ask for 6.5 seconds and Mirage bills you for 12 at $2.10 while fal bills 6.5 at $1.95. Anywhere under 7 seconds of output in that block, the reseller wins on price.
That window is narrow and it is a rounding artifact, not a strategy. But it is the cleanest possible demonstration of the point: a rate comparison that ignores the increment can be wrong in either direction, and the direction it is wrong in depends on a number you have not generated yet.
The rule worth keeping: a price is a rate and a unit, and a vendor who publishes only one of the two has not published a price.

The two surfaces also disagree about the model itself
Prices are not the only thing that diverges when a model is sold in two places. The specs do too, and the disagreements are load-bearing.
Sixty seconds, or one hundred and eighty
Ask each surface how long a video can be and you get a different answer:
- fal's schema: generated speech must be 180 seconds or shorter.
- Mirage's own docs: driving audio is 1 to 60 seconds per request, and for longer videos developers can generate multiple segments and join them together.
Both statements are current. They are probably both true of their own surface, describing different arms and different constraints. What you cannot do is quote one maximum as "the model's maximum."
The budgeting consequence is concrete, and it is not only about money:
- Three minutes of talking head at fal's rate is $54, and the schema says you can ask for it in one request.
- Three minutes through Mirage is at least three requests of 60 seconds each, 30 six-second blocks in total, or $31.50.
- The $22.50 difference comes with a workload transfer. The cheaper path is the one where you write the segmentation, the joining and the retry logic.
That is a real trade, and it is the honest form of the whole comparison. You are not choosing between two prices. You are choosing between a price and a price plus some of your own engineering.
The resolution that only one page publishes
Mirage's docs say output is a 720p MP4 in 9:16 or 16:9, with most generations completing in under two minutes. fal publishes no resolution at all for these endpoints.
Line the two surfaces up field by field and the asymmetry is easy to see:
- Rate: published by both, and different.
- Billing increment: published by Mirage, absent on fal.
- Resolution: published by Mirage, absent on fal.
- Aspect ratio: stated by Mirage, implied on fal by which avatar name you pick.
- Maximum length: published by both, and in conflict.
- Supported languages: published by neither.
So if you are comparing this model against something else on picture quality, the 720p figure is a vendor-documented number and you should say so when you cite it. A fal-only citation cannot support it, and "the reseller did not mention resolution" is not evidence that the resolution is different. Absence on a reseller's page is absence of a statement, nothing more.
What neither page establishes
Languages. Neither surface publishes a supported-language list, and there is no language parameter to set: fal's complete input contract for the text-to-video arm is four fields, script, avatar, video_reference_url and audio_reference_url. That is the entire schema.
Write that down as not established rather than assuming English-only or assuming multilingual. Three things belong in that column, and keeping them there is the difference between an evaluation and a wish list:
- Languages. No list, no parameter, no statement either way on either surface.
- Any independent evaluation. The comparative claims that exist are the vendor's own two uncited captions, and Mirage's own press tracker lists no Avatar X coverage to set beside them.
- The Captions app features. The widely repeated line about ten seconds of video being enough to make your twin belongs to the consumer app, not to this API. Carrying it across is how a spec sheet acquires a fact nobody published.
What to do with a superlative nobody measured
Both listings carry the same sentence, and it is worth being precise about where it comes from.
Where "industry-leading" comes from
fal's catalog describes the endpoint as delivering industry-leading identity preservation and expressivity in AI video. That is not fal's assessment. It is the shortDescription the model lab supplies to the catalog, which is why it appears word for word on both arms and reads like marketing, because it is.
The vendor's own comparative claims sit in the launch post as two uncited chart captions: that Mirage Avatar X is more expressive than other models, and that competitive models show less organic emotion. No competitor is named, no methodology is given, no numbers are attached.
What Mirage publishes when it has evidence
This is the fair framing, and it cuts both ways. Mirage does publish evaluated research when it has some. Its predecessor audio-to-video model has a full paper, "Seeing Voices: Generating A-Roll Video from Audio with Mirage", on arXiv from June 2025. That is a real artifact with a method section.
Avatar X got a marketing post instead. And as of August 11, 2026, Mirage's own press page, which curates dozens of items back to 2023 across TechCrunch, Forbes, Fast Company and others, lists no Avatar X coverage at all. A company that tracks its own coverage that carefully would list it if it existed.
None of that makes the model bad. It makes the superlative unverified, which is a different thing, and it means the only evaluation that exists for your use case is the one you run.
The three questions to ask instead
For any avatar model you are about to put in a production pipeline:
- Does the same face survive across a set of variants, not just one flattering demo? Identity drift is what breaks creative testing, because ten variants only teach you something if they look like the same person.
- Does the voice hold up on a phone speaker at low volume, which is how it will actually be heard?
- What does one usable output cost, counting the retries, rather than what does one second cost?
That third question is why the increment matters more than the rate. It is the same reason people comparing HeyGen alternatives or Synthesia alternatives on headline seat price end up surprised: the published number and the number on the invoice are related, but they are not the same number.
How Novoads solves the price you did not read
Novoads does not run Mirage Avatar X. There is no Mirage or Captions model in the platform, and this post is a read on an adjacent listing rather than a product announcement.
What is relevant is the shape of the problem, because it is one we made a deliberate choice about. Per-second billing with a hidden increment asks you to predict the length of a thing that has not been generated yet, then reconstruct the invoice afterwards. Novoads prices a render in credits, quoted against the model and settings before you run it, so the cost of the output is a number you read rather than a number you derive. That is also why our own model pricing pages state the schedule per clip rather than per second. When a tool quotes in credits instead of seconds, the same arithmetic is waiting for you with a private unit in the middle of it.
The rest of the workflow is the ordinary one:
- Upload a product image.
- Write the script, or have it generated from the product.
- Pick a model and read what that render will cost before you run it.
- Produce the UGC-style ads you are going to test against each other.
If you want to try it, Novoads runs a $1 trial. $1 for 3 days of access, cancel whenever you want.

Distribution is a price, not a convenience
The convenient door is rarely the cheap one, and that is not a scandal. A reseller carries the integration, the uptime, the unified billing across a dozen models, and a single key that already works in your codebase. Those are worth money, and 71% per second may well be a fair price for them on a low-volume project where engineering time costs more than inference.
What is not defensible is not knowing. The model lab's own pricing page is one search away, it publishes both the rate and the increment, and reading it before wiring up the reseller is the entire discipline. It generalizes past this one model: whenever a thing you buy is available from the people who made it and from someone standing next to them, check the maker's page and check what unit each side counts in. Two vendors quoting per second are not necessarily quoting the same second.
Concretely, for this model on this pair of surfaces:
Buy direct from Mirage when:
- Your volume is high enough that a 71% list gap outweighs the integration time you would save.
- Your outputs land near 6-second boundaries, where the gap is at its widest and the rounding costs you nothing.
- You want the resolution, the aspect ratios and the provenance behavior documented at all, which today only the vendor's own docs do.
- You are already going to write segment-and-join code, because your pieces run past 60 seconds anyway.
Buy through fal when:
- This model is one of several you already call behind a single key, and one integration is the whole point.
- Your clips sit in the awkward part of a 6-second block, where per-second billing is genuinely competitive.
- The project is small enough that a day of engineering costs more than the markup ever will.
- You are still deciding whether this model belongs in the pipeline, and the cheapest thing to buy is the ability to stop.
Buy the convenience if you want it. Just buy it knowing what it costs, and knowing what a UGC creator or an avatar built from one photo would have cost you instead.
Frequently Asked Questions
How much does Mirage Avatar X cost?
It depends which door you buy it through, and both prices are current as of August 11, 2026. Mirage's own API pricing page lists Avatar X at $0.175 per second, billed on the length of the generated output rounded up to the nearest 6-second increment. fal, which resells the model as a partner endpoint, states that your request will cost $0.3 per second and publishes no minimum increment. The per-second list rate on fal is about 71% higher.
Is the model on fal the same model Mirage sells?
Yes. fal's own catalog data labels both Avatar X endpoints with model lab Mirage and model family Avatar X, and the description fal publishes is Mirage's own marketing copy word for word. It is a partner listing of the vendor's API, not a lookalike or a fine-tune.
Does the 71% gap mean I pay 71% more on fal?
Not always, and this is the part worth checking before you wire anything up. Mirage bills in 6-second blocks, so a 6.5-second output is billed as 12 seconds there. fal bills per second with no stated minimum. The gap is 71% when your output lands exactly on a 6-second boundary, falls to about 43% on a 15-second clip, and for outputs just past 6 seconds fal is briefly the cheaper of the two.
What is Mirage, and what happened to Captions?
They are the same company. The footer on both mirage.app and captions.ai reads NOCAP, Inc. d/b/a Mirage, and TechCrunch covered the rename in September 2025. The Captions app still exists as a product; Mirage is the research and model side of the same business, which is why the Avatar X API docs and pricing page live on captions.ai.
How long can an Avatar X video be?
State which surface you mean, because they conflict. fal's schema says generated speech must be 180 seconds or shorter. Mirage's own API documentation says driving audio is 1 to 60 seconds per request and that for longer videos developers can generate multiple segments and join them together. If you are budgeting a three-minute piece, that difference is the difference between one request and at least three.
What languages does Avatar X support?
Not established. Neither surface publishes a supported-language list, and fal's complete input schema for the text-to-video arm has exactly four fields: script, avatar, video reference and audio reference. There is no language parameter to set. If a specific language matters to your campaign, treat it as an open question to test rather than a documented capability.
Key Takeaways
- Mirage lists Avatar X at $0.175 per second on its own API pricing page, billed on the generated output length rounded up to the nearest 6-second increment. fal lists the same model at $0.3 per second, with no stated minimum increment.
- That is a 71% higher list rate per second on the reseller. It is a list-rate gap, not automatically what you pay, because the two surfaces do not bill in the same unit.
- The increment is the part people skip. Because Mirage rounds up to 6-second blocks, the real gap is 71% only when your output lands on a block boundary, drops to roughly 43% on a 15-second clip, and briefly inverts in favor of fal just past 6 seconds.
- Mirage is NOCAP, Inc., the company that used to trade as Captions. captions.ai and mirage.app are the same firm, which is why the model's pricing page still sits on the old domain.
- The two surfaces also disagree about the model itself. fal says generated speech must be 180 seconds or shorter; Mirage's own docs say driving audio is 1 to 60 seconds per request and that longer videos need multiple joined segments.


