LEMON AI/RESOURCES

Creative testing across countries

Two Thirds of a Creative Ranking Survives the Border

We ranked the same AppLovin creatives in two countries at once. 71% of the cost-per-install signal crosses; the local read still beats the global pool.

Creative testing across countries

We ranked the same AppLovin creatives in two countries at once. 71% of the cost-per-install signal crosses; the local read still beats the global pool.

A creative that ranks in the top quarter on cost per install in one country is in the top quarter of the next country 39.6% of the time. Inside its own country, on the same amount of data, it is there 46.8% of the time. Chance is 22.5%. So roughly two thirds of whatever the ranking knows crosses the border, 71% with a confidence interval of 55 to 79%, and more than 80% of it crosses between tier-one markets. That is the first half of the answer. The second half is that a country's own ranking, built from a fraction of the installs, still predicts that country better than the pooled global ranking does. Both are true at once, and most of the advice on this decision ignores one of them.

The measurement covers 964 AppLovin creative sets delivering in 48 countries across 42 apps, June 22 to August 30, 2026, with the country dimension taken from AppLovin's own advertiser report and nothing allocated by us.

Top-quarter survival of a creative cost-per-install ranking. Across all 2,575 country pairs: 46.8 percent inside one country, 39.6 percent across a border, 22.5 percent by chance. Between tier-one markets: 51.7, 46.8 and 22.1. Tier one to the rest: 48.3, 39.3 and 22.4. Outside tier one: 44.7, 38.6 and 22.6.

The advice assumes the ranking travels

Two versions of the same decision come up whenever a team runs one app in many markets. The first is to test creatives where impressions are cheap and scale the winners where they are expensive. A seller on r/FacebookAds describes it exactly: test in Poland and Hungary first, then expand the winning ads to the US and Canada. The second is to launch a new market on the home market's winners. RocketShip HQ's multi-country guide puts a number on it: "about 40% of top-performing creatives in the US also perform well in other Tier 1 markets with proper localization", stated as experience, with no sample, no method and no line for what a coin flip would produce.

The opposite school says the ranking does not travel and each market needs its own creative. Bruin's analysis template asks whether the top five creatives show CPI variance above 40% across tier-one geos, and if they do, recommends geo-specific creative rather than one global set. Liftoff's market-entry post from 2018 has Hungarian users preferring a profile-style ad where neighbouring markets preferred the glossy one.

None of these positions is measured against the one thing that would settle it: how well the same ranking reproduces when nothing changes except the border. A ranking that only reproduces at 47% inside one country cannot be expected to reproduce at 90% across one. The question is not whether the cross-border number is high. It is how much of the within-country number survives.

On AppLovin the question has a sharper edge, because the buyer rarely chooses which country sees which creative. AppLovin's own scaling guidance says to "select all countries that you actively support" and to set "a single, global budget for optimal spend allocation towards the best outcomes across all countries". The engine spreads the creative set across markets; the team reads a blended row afterwards and decides what it means for the market they care about.

Rank the same creatives in two countries at once

The test is the one we used to measure how much spend a creative ranking needs before it holds, with a border added.

Take one comparison scope: one AppLovin advertiser account, one app, one operating system, one fortnight. Split the fortnight into alternating days, so every creative set has an odd-day half and an even-day half in every country it delivered in. Then, for every pair of countries that share at least eight creative sets:

  1. Rank the shared creatives on cost per install in country A's odd days, and again in country A's even days. The share of the top quarter that appears in both is the within-country survival.
  2. Rank them in country A's odd days and in country B's even days. The share of the top quarter that appears in both is the cross-border survival.
  3. Permute country B's ranking and repeat the second step sixty times. That is the chance line.
  4. Report the share of signal retained: cross minus chance, divided by within minus chance.

Both comparisons use halves of identical size from the same fortnight, so the border is the only thing that changes between steps one and two. The chance line is measured rather than assumed because a top quarter of eleven creatives is two creatives, and two out of eleven is not 25%. Ties are broken at random and every figure is averaged over twenty-five draws. Confidence intervals bootstrap over comparison scopes, because creatives inside one app are not independent of each other.

A creative set and country qualify when each half carries at least $25 and three installs. The floors are deliberately low so that the panel is wide; the sensitivity checks below raise them. The tier-one label is a convention we chose before running anything: United States, United Kingdom, Canada, Australia, Germany, France, Japan and South Korea. Nothing about the result depends on it beyond the split.

The panel behind the comparisons is 6 AppLovin advertiser accounts, 248 apps, 21,241 creative sets and 236 countries, $16.3 million of spend and 25.1 million installs. After the floors, 964 creative sets in 48 countries across 42 apps and 4 accounts take part in at least one comparison, carrying $4.3 million of spend, 1.85 million installs and 128 million impressions. A typical creative-country half is $117 and 21 installs. That produces 2,575 country-pair comparisons from 89 app-platform-fortnights. Two accounts' feeds end on August 10; their partial fortnight is excluded rather than half-counted.

What crossed the border

Country pairComparisonsInside one countryAcross the borderChanceSignal retained
All pairs2,57546.8% (38.6 to 51.0)39.6% (33.3 to 44.3)22.5%71% (55 to 79)
Tier one to tier one24051.7% (44.8 to 57.4)46.8% (39.0 to 53.5)22.1%83% (68 to 95)
Tier one to the rest1,00848.3% (41.0 to 53.3)39.3% (32.5 to 44.6)22.4%65% (43 to 80)
Outside tier one1,32744.7% (33.9 to 49.9)38.6% (29.9 to 43.3)22.6%73% (57 to 80)
Any pair including the US38255.5% (50.0 to 60.4)43.0% (39.0 to 46.9)22.6%62% (54 to 71)

Ranges are 95% confidence intervals. Signal retained is (across minus chance) divided by (inside minus chance).

Read the first row as the plain answer. A cost-per-install ranking that would reproduce at 46.8% if you simply ran it again in the same country reproduces at 39.6% if you run it again in a different one. The border costs seven points out of the twenty-four the ranking had above chance.

The tier-one row is the strongest transfer in the panel. Between two expensive markets the ranking keeps 83% of its signal, and the confidence interval does not reach below two thirds. Crossing from tier one into the rest of the world keeps 65%, and that interval is wide enough to include a half. The US rows are the cleanest scopes in the panel, with the highest within-country reproduction at 55.5%, and also the largest absolute drop, 12.5 points, because the US ranking had the most to lose.

The single best creative tells the same story from the top of the list. Inside one country, the creative that ranked first on odd days lands at the 25th percentile of the even-day ranking, in the median comparison. Across a border it lands at the 31st. For US pairs the numbers are 12.5 and 27.5: the American winner is still in the top third abroad, and is no longer near the top.

RocketShip's 40% is not far from the tier-one row. What it lacks is the two lines around it. Forty percent sounds like a coin flip until you know that chance is 22% and the same country would have given you 52%.

Raising the floors does not change the shape, and it does thin the panel. At $100 and ten installs per half, 186 comparisons from 16 scopes reproduce at 49.2% inside a country and 38.2% across a border, against 23.1% chance, retaining 58% with a wide interval of 34 to 87%. Tier-one pairs at that floor keep 86% of a much stronger 69.4% within-country signal; tier one to the rest keeps 41%, with an interval from 21 to 82%. Requiring at least twelve shared creatives instead of eight, 1,179 comparisons reproduce at 49.6% inside and 42.8% across, retaining 74% (57 to 93). The cross-border figure moves with the within-country figure, which is what a genuine transfer of signal looks like and what a floor artefact would not.

The price gap barely matters

The intuition behind cheap-geo testing is that a cheap market is a different market, so the further apart two countries are in price, the less one says about the other. The panel supports only a weak version of that.

CPI ratio between the two countriesComparisonsInside one countryAcross the borderSignal retained
Under 1.3x56045.7%39.0%72%
1.3x to 1.6x53146.0%40.7%78%
1.6x to 2.7x79346.9%40.0%71%
Over 2.7x69148.1%38.9%64%

The ratio is computed on the same creatives in the same fortnight, so it is the actual price gap the team faced, not a tier table. Transfer falls only in the last bin, where one country costs almost three times the other per install, and even there it keeps 64% of the signal. A test market that is a little cheaper is not measurably worse as a proxy than one that is a little dearer. A test market that is a lot cheaper is somewhat worse.

The local read beats the global pool

If most of the signal travels, the obvious shortcut is to skip the local ranking altogether and rank on the pooled global row, which has vastly more installs behind it. We tested that shortcut directly. For every country with at least eight qualifying creatives in a scope, we predicted its even-day ranking three ways: from its own odd days, from the pooled odd days of every other country the same creatives ran in, and from the full fortnight of every other country.

Predictor of the country's held-out rankingMedian installs behind itTop-quarter survival
The country's own other half85050.1% (47.4 to 52.8)
Every other country, same half11,21345.6%
Every other country, full fortnight22,42546.5% (43.3 to 49.6)
Chance22.7%

640 country-fortnights. Installs are medians across those country-fortnights and are summed over the creatives in the comparison.

The pooled global ranking predicts a country's next reading worse than that country's own previous reading does, with twenty-six times the installs behind it. The paired difference is 3.6 points in favour of the local read, with a 95% interval of 0.9 to 7.0, and the local read wins or ties in 74% of country-fortnights. Split by tier the pattern holds: 50.4% against 47.6% in tier-one countries, 49.7% against 45.0% outside them, where the pool had seventy-three times more installs than the local half.

That is the second half of the answer, and it is the half the pooling shortcut gets wrong. Geography carries information about the creative ranking that no amount of other-country data recovers. The global row is a good prior at 46.5%, more than twice chance. It is not a substitute for the country's own reading once that reading exists.

Revenue does not travel because it does not reproduce

On day-7 ROAS, from AppLovin's cohort-mode report, the ranking is at chance in both directions: 25.2% inside one country, 23.7% across a border, against a chance line of 22.9%, on 143 comparisons drawn from a cohort panel of 8,332 creative sets, $8.0 million of cohort-report spend and $1.1 million of day-7 revenue. Raising the floor to $100 per half does not move it: 22.3% inside, 20.2% across. The creative that ranked first on revenue lands near the median of the other ranking, at the 44th percentile, whether or not a border is in the way.

This is the same result the spend-threshold measurement found without a border, and it has the same cause. A creative's day-7 revenue in one country in one week is, for most creatives, one or two purchases. Nothing built on that reproduces, so nothing built on it can transfer. A team porting its "best ROAS creative" into a new market is porting noise with a confident label on it.

What to do with a winner from another market

Launch on the borrowed ranking, on the delivery metric. A cost-per-install ranking from another country is worth 39.6% against 22.5% chance, and a tier-one ranking is worth 46.8% in another tier-one market. That is a real prior and it costs nothing. Use it to choose which creatives go live first in the new market.

Replace it with the local ranking as soon as the local ranking clears your own reproduction floor. The local read beat the pool at a median of 850 installs per half-fortnight, and 358 outside tier one. Your floor is the point on your own reproduction curve where survival separates from chance, and the procedure above gives it to you in one query on a creative-set-by-country export.

Treat a cheap test market as a discount, not a proxy. Testing in a market that costs a third of your target market keeps about 64% of the signal on cost per install. It is cheaper evidence about the target market, and it is weaker evidence. Budget for confirming the top of the list locally before scaling.

Do not port a revenue ranking anywhere. It did not reproduce at home. Rank on cost per install or installs per mille for the creative decision, and judge revenue at the campaign or country grain, where there are enough purchases to rank.

Run the test on your own account before you trust any of these numbers. Pick a fortnight, split it into odd and even days, and for each pair of countries that share eight or more creatives, count how many of the top quarter in one country's odd days are in the top quarter of the other country's even days. Permute one side a few hundred times for your chance line. If the cross-border figure sits on that line, the ranking you are about to port is not carrying a decision, whatever it looked like in the dashboard.

When a borrowed winner does hold locally, the next question is why, and that is a test rather than a tag. The creative test brief turns a scoped finding like this one into a single controlled change with predeclared result branches.

What this does not establish

The network chose the exposure in both countries. AppLovin decided which impressions each creative set got in country A and in country B. A creative that wins in A may have been handed the easier inventory in B because it won in A, or the same inventory because the engine judges the creative globally. The measurement cannot separate the creative from the allocation. What transfers here is the reported ranking, which is the object teams act on, not the creative's causal effect in the new market.

This is one network, at creative-set grain, over ten weeks. Asset grain was rejected because Lemon reconstructs asset-level country splits by allocating the creative set's reported distribution, which would have made the test circular. Google's asset feed carries no country dimension. Meta was not measured. An app whose creatives are heavily localized, or a team running a handful of creatives at far higher spend per country, will see different numbers.

Nothing here says whether localization pays. The panel contains whatever mix of translated, subtitled and untouched creative sets these accounts ran, and does not know which is which. The result is that the ranking of the creatives as delivered travels; it is silent on whether a localized version would have ranked higher.

The tier split is a convention. Eight countries were labelled tier one before the analysis ran. Other lists exist, and the price-gap sweep, which uses the measured cost ratio instead, is the less arbitrary version of the same cut.

The panel is US-heavy. The United States is 50% of the qualified spend and tier-one countries are 69% of it. Confidence intervals for the outside-tier-one rows are wide because those comparisons come from 28 scopes.

Where Lemon fits

Lemon's asset table filters by country and marks which figures are network-reported and which are reconstructed. On AppLovin, the country split at asset grain is reconstructed from the creative set's reported country distribution, which is exactly why this article measured creative sets and not assets; the creative-set rows carry the country the network reported. That makes the two readings this article compares, the country's own ranking and the pooled one, available side by side, with their provenance visible. It does not make the cross-border transfer any higher than the panel says, and it does not tell you whether a creative that wins in the new market does so because of the creative or because of the delivery it was given there. The decision guide to mobile app creative analytics covers where each reading belongs in the weekly cadence, and half of creative decline flags are country mix covers the other direction of the same dimension, when the mix moves and the creative did not.

Primary sources

NEXT STEP

Bring us the question your dashboards can't answer.

See your creatives and cohorts in one place, with the answers your ad networks leave out.

Book a demo

LEMON AI

Bring us your growth question.