Methodology
How Lemon turns source data into decisions
How Lemon AI reconciles ad-network data, attributes outcomes, aggregates ratios, forecasts cohort revenue, backtests models, and communicates limitations.
Core rule: Lemon distinguishes network-reported values, joined advertiser outcomes, reconstructed creative outcomes, and model predictions. A displayed number must retain enough provenance to explain which class it belongs to.
Scope and purpose
Lemon AI supports two related decision systems. Creative Intelligence joins ad assets to delivery and business outcomes so a mobile growth team can decide what to scale, stop, or make next. Cohort Prediction estimates future revenue and ROAS before every acquired cohort reaches its full maturity horizon.
The methodology is designed for acquisition decisions, not statutory accounting. Network, MMP, warehouse, billing, and finance systems can use different correction schedules and definitions. Lemon preserves those distinctions instead of asserting that one dashboard is the universal financial ledger.
Source systems and grain
Each integration begins with a source contract. The contract records authentication scope, available entities, stable identifiers, reporting timezone, currency, supported dimensions, metric definitions, latency, correction behavior, and API limits. Data is ingested at the most granular trustworthy level the source exposes; a dashboard request cannot create a dimension that the source never supplied.
| Source class | Typical contribution | Important limitation |
|---|---|---|
| Ad network | Asset identity, spend, impressions, clicks and campaign structure | Downstream outcomes may not exist at the same asset grain |
| MMP | Attributed installs, events, purchases and revenue | Attribution windows and correction timing differ |
| Advertiser warehouse | Validated product events, subscriptions or internal revenue | Event identity and availability vary by customer |
| Lemon model | Future cohort revenue and derived predicted ROAS | Predictions are estimates, not observed revenue |
Identity and deduplication
Network object IDs are retained separately from human-readable names. Names can change or be reused; they are not treated as universal keys. Creative media can also appear in more than one campaign or upload. Lemon retains the source asset, ad, campaign, account, app, platform, country, and observation period required to disambiguate those appearances.
Source rows are deduplicated using their natural identifiers and reporting scope. Reruns replace or merge the same logical observation rather than adding a second copy. If a source restates recent data, the retained observation is updated and downstream aggregates are recomputed.
Timezone, currency, and reporting windows
Dates are normalized deliberately, not inferred from the machine running the job. API requests and stored observations retain the source timezone or documented UTC boundary. When two sources use different reporting days, reconciliation occurs only after converting them to a declared comparison boundary.
Currency values retain their source currency and any normalized currency separately. Exchange-rate source and effective date must be known before values are combined. A report does not sum monetary values across currencies without normalization.
Attribution and outcome alignment
An install, purchase, or revenue event belongs to a creative only when the available source relationships support that assignment. Click-through and view-through attribution windows, re-engagement rules, event deduplication, and privacy-related aggregation can all change the relationship between a network delivery row and an advertiser outcome.
Lemon does not silently force unmatched outcomes into creative rows. Coverage and unmatched totals remain observable. This is particularly important when the advertiser total is known but the network does not expose a compatible creative breakdown.
Creative-level reconstruction
When a downstream outcome is not reported directly at asset level, Lemon can reconstruct an asset-level estimate within a reconciled scope. The scope supplies a constraining observed total. Only eligible creative rows inside that same app, campaign, platform, country, and time scope can receive an allocation.
The result is marked as reconstructed rather than network-reported. The allocation must be reproducible, preserve the relevant constraining total within documented rounding tolerance, and leave unsupported value unmatched. Reconstruction does not convert an estimate into a source fact.
Additive and non-additive metrics
Spend, impressions, clicks, installs, purchases, and compatible revenue amounts are additive inside a valid scope. CTR, CPI, CPA, conversion rate, ARPU, and ROAS are ratios and are not additive. Lemon aggregates their underlying numerators and denominators, then recomputes the ratio for the requested group.
For example, country-level ROAS is total aligned revenue divided by total aligned spend for that country. It is not the arithmetic average of its creative-level ROAS values. This rule prevents small rows from receiving the same weight as high-spend rows.
Cohort forecasts
A cohort is defined by a consistent acquisition date and scope. At a given observation age, Lemon combines the behavior available so far with a trained model to estimate revenue at a named future horizon. The display distinguishes actual revenue, predicted remaining revenue, forecast total, and predicted ROAS.
Forecast inputs must be available consistently at inference time. Data-quality checks cover missing features, event-schema changes, unexpected ranges, traffic-mix shifts, delayed data, and training-serving skew. A prediction can be withheld when required input coverage is not sufficient.
Backtesting and error
Backtests use time-based evaluation. Training data must precede the evaluated acquisition period, and the prediction uses only observations that would have been available at the intended decision age. Matured revenue at the target horizon supplies the comparison.
Evaluation includes absolute error, weighted error, bias, and MAPE where actual values make MAPE interpretable. Error is inspected by horizon, app, platform, country, media source, spend scale, and time period. No single aggregate percentage is treated as proof that every cohort or market has the same reliability.
Freshness and revisions
Every reporting surface should expose the latest successful source update. Recent dates may change as networks or MMPs process delayed events and corrections. Historical values can also change when the customer corrects an event definition or source mapping. Material logic changes require a before-and-after validation against known healthy data.
Limitations
- Source APIs can be delayed, incomplete, rate-limited, or retrospectively corrected.
- Privacy frameworks can reduce user- or asset-level observability.
- Attribution is a defined measurement model, not direct observation of causality.
- Reconstructed creative outcomes inherit the assumptions and coverage of their allocation.
- Forecasts can degrade after product, traffic, pricing, market, or monetization changes.
- Small cohorts and near-zero actual values produce unstable percentage errors.
- Cross-platform and cross-network totals are comparable only after definition alignment.
How to use the output
Select the metric, scope, maturity horizon, and minimum evidence before ranking assets or cohorts. Inspect source freshness and coverage. Compare like-for-like segments. Use a forecast as a decision input with known error, not as booked revenue. Record the budget or creative action and evaluate it in the next review cycle.
Continue with the guides to mobile app creative analytics, AppLovin creative reporting, and cohort revenue forecasting.