AppLovin creative testing
AppLovin Creative Testing When the Engine Allocates Budget
AppLovin decides which creative spends your money, then reports twelve delivery columns and no outcomes. What a creative comparison can and cannot prove.
On AppLovin the engine decides how much each creative spends, so a spend ranking is an output of that decision and not a measurement of creative quality. The asset report gives you twelve delivery columns and no outcome columns, which means AppLovin’s own data cannot tell you whether the asset it funded was the asset that earned.
This page sets out what a creative comparison on AppLovin can legitimately prove, what it cannot, and the guards that keep a decision honest when you did not control the allocation.
AXON is now AppLovin Ads
If you are searching for AXON, you are searching for a name AppLovin has stopped using. The advertising platform was rebranded to AppLovin Ads on June 30, 2026, and the self-serve product opened to all advertisers rather than a referral programme.
Verified on August 15, 2026: AppLovin’s own advertiser page refers to AppLovin Ads throughout and describes the optimiser only as “AppLovin’s AI engine”. The documentation host support.axon.ai returns HTTP 301 to support.applovin.com on every path we tested.
Nothing underneath changed. The engine still buys, still allocates, and still declines to explain itself. Practitioners will keep saying AXON for years, which is fine. Just do not expect current documentation to be filed under that word.
What the engine reports about its own decisions
Nothing.
That is the whole finding, and it is worth being precise about it. As documented on August 15, 2026, AppLovin’s Asset Reporting API exposes exactly these columns:
asset_id, asset_name, asset_url, campaign, campaign_id, campaign_package_name, clicks, cost, creative_set, creative_set_id, ctr, impressions
Two endpoints serve them, and they do not behave the same way:
| Endpoint | Date handling | Ceiling |
|---|---|---|
/assetReport |
Presets only: yesterday, last_7d, last_month |
No custom range |
/assetAnalyticsReport |
Custom range, YYYY-MM-DD |
45 day maximum window |
Read the column list again for what is missing. No installs. No revenue. No ROAS, CPI, or CPA. No conversion of any kind at asset grain.
And nothing about the engine. No exploration or exploitation state is documented, nor a learning-phase marker, a per-asset budget allocation, a reason code, or a confidence value. The system that moved your money publishes no field describing why it moved.
That opacity is not accidental, and AppLovin’s leadership has been direct about it. Speaking to AdExchanger in February 2024, chief executive Adam Foroughi said of the model, “We can’t see into a black-box algorithm.” The campaign scaling guidance tells advertisers to upload and test new assets frequently and to diversify formats. It documents no learning period and no ramp timeline.
So the instruction is to keep feeding creatives, and the reporting surface will not tell you which ones worked.
Why this is not a test
A creative test assumes you controlled the exposure. On AppLovin you did not. The engine chooses how much each asset spends, and it chooses continuously.
That produces three confounds you cannot remove with the asset report alone.
Allocation is the treatment. Assets the engine favours receive more spend, more impressions, and therefore more mature outcome data. When you compare them to assets it starved, you are comparing two different levels of statistical power alongside two different creatives.
Selection runs both ways. The engine allocates on its own predicted outcome. An asset that looks strong early attracts budget, which produces the volume that makes it look strong in your report. Rank and cause are entangled from the first day.
The report stops at delivery. Impressions, clicks, cost, and CTR describe what the engine did. Whether it was right is a question about installs and revenue, and those are not in the response.
None of this means the engine is allocating badly. It means the asset report cannot answer the question you are asking it.
What our own data shows
We measured this across our production AppLovin accounts, using $32.4 million of AppLovin spend on 742 creative assets, with install-cohort D7 outcomes for cohorts from June 17 to August 3, 2026. Every asset carried at least $1,000 of spend and 50,000 impressions, and every app contributed at least five qualifying assets so that ranking meant something. Outcomes at asset grain are reconstructed, because AppLovin does not report them there. ROAS is recomputed from summed revenue over summed spend, never averaged.
Across that $32.4 million, ranked by D7 ROAS against the assets running for the same app, the highest-spending asset landed at the 50th percentile. The middle. The interquartile range ran from the 40th to the 68th percentile.
The obvious objection is that a low-volume asset can win a ratio on luck, which would manufacture this result. It does not. Raising the floors moves the number very little:
| Minimum per asset | Assets | Spend | Top spender’s median rank |
|---|---|---|---|
| $100 and 10k impressions | 1,467 | $33,201,700 | 53.6th percentile |
| $500 and 25k impressions | 1,042 | $32,887,116 | 45.0th percentile |
| $1,000 and 50k impressions | 742 | $32,412,650 | 50.0th percentile |
| $5,000 and 100k impressions | 319 | $31,178,147 | 60.0th percentile |
Across every floor the biggest spender sits between the 45th and 60th percentile of return. It never approaches the top.
Spend is also extremely concentrated. The top 10% of assets by spend took 88.4% of all spend, and the top 25% took 95.1%. Most of the budget rides on a handful of creatives whose position in the return ranking is close to average.
The engine is not allocating at random, and we are not claiming it is. Spend-weighted D7 ROAS across these assets was 0.072, while the median qualifying asset returned 0.0115. Weighted by where the money actually went, the portfolio performs several times better than the typical asset. Allocation carries real signal.
Both things are true, and the useful conclusion sits between them. The engine puts money in a better than average part of the distribution and does not find the best asset in it. That is exactly what you would expect from a system optimising a predicted outcome under uncertainty, and it is why the spend column is a poor proxy for creative quality.
These figures describe our client base, not the market. They are one vendor’s accounts over seven weeks in mid 2026, at D7 only, and ranks can still move by D30. Delivery dates and install-cohort dates carry different semantics, so aligning them over one window is an approximation.
What you can and cannot conclude
| Question | Evidence it needs | In the asset report? | Do this instead |
|---|---|---|---|
| Which asset did the engine back? | Cost by asset | Yes | Read it directly. This is the one thing the report answers cleanly. |
| Which asset earned the most? | Revenue by asset | No | Reconstruct from your MMP or warehouse at asset grain |
| Which asset earned most per dollar? | Revenue and spend, aligned | Partly | Recompute the ratio from aligned components; never average ROAS |
| Is asset A better than asset B? | Comparable exposure | No | Treat as a ranking hypothesis, then confirm on matured cohorts |
| Did this creative fatigue? | Rate series with volume floors | Delivery only | Run the fatigue measurement protocol |
| Why did the engine choose it? | Allocation state | No | Not answerable. Stop trying to infer it |
The last row is the one teams waste the most time on. There is no field, so every explanation of the engine’s behaviour is a story fitted after the fact.
Guards that survive an allocated test
Four rules keep a conclusion defensible when you did not control delivery.
Set a volume floor before reading, not after. A ranking over assets that never got enough delivery is noise. Our own floors above are a starting point rather than a universal rule; tune them to the variance in your own account.
Match observation windows. Compare assets over the same calendar range. The 45 day ceiling on /assetAnalyticsReport means longer comparisons must be stitched from multiple pulls, and a stitched window is easy to misalign.
Gate on outcome maturity. D7 ROAS is only readable once every cohort in the window has had seven days. Comparing a mature asset against one still accumulating manufactures a gap that closes on its own. Which horizon is even available depends on the network, and the attribution window comparison sets out what AppLovin anchors its outcomes to and when the figure stops moving.
Keep provenance visible. Any outcome you attach to an AppLovin asset was reconstructed. Label it. The asset-level creative ROAS audit covers the provenance, scope, maturity, and aggregation checks a number should pass before it changes a budget.
Where this leaves the work
Feed the engine variety, because it can only choose among what you give it. Then build the evidence layer it does not provide: join asset-grain delivery to attributed outcomes, keep an unallocated bucket for what will not match, and rank on reconstructed return rather than on spend.
That work is the same whether you do it in a warehouse or in a product. Lemon AI’s creative analytics puts AppLovin’s reported delivery next to reconstructed installs, purchases, revenue, and ROAS for the same asset, with recovered values visually distinct from network-reported ones so the provenance stays legible. The reconciliation rules and limitations are published in full.
For how the AppLovin reporting surfaces fit together once you have decided what to measure, see AppLovin creative reporting.