The first question is not what happened. It is what still exists
In most commercial disputes counsel begins with the conduct and then goes looking for the documents. In a digital marketing dispute that order fails, because the documents are on a timer. The advertising and analytics record is generated by platforms, held by platforms, and deleted by platforms on published schedules that run whether or not a complaint has been filed, whether or not a hold letter has gone out, and whether or not anyone has yet realized the case turns on it.
So the first question I ask on a new matter is not what the parties did. It is what survives, what expired, and when. That answer shapes everything downstream — the theory, the discovery plan, the damages model, and whether an expert can say anything useful at all. A claim that would have been straightforward at week three is often unprovable at month nine, not because the facts changed but because the record aged out.
This is the orientation page for that question. It describes where the record lives, how long it lives, which parts of it are observations and which are estimates, and what none of it settles.
Who actually holds the record, and why it is rarely your client
An advertiser looking at a campaign dashboard is looking at a rendered view of somebody else's database. The underlying rows sit with Google, Meta, Amazon, an ad server, an email service provider, or an affiliate network. The advertiser has credentials and an export button. It does not have the data.
That distinction matters twice. It matters for obtaining the evidence, because the fastest route is almost always the party's own self-service export rather than a subpoena to the platform — third-party process here is slow, contested, and of uncertain scope, and the Stored Communications Act's treatment of advertising business records is not a settled question. And it matters for preservation, because a party with credential access and a practical ability to export has been held to have Rule 34 control over third-party-hosted data. In Brown v. Tellermate Holdings (S.D. Ohio 2014) the court rejected a defendant's claim that it neither possessed nor controlled records held in a vendor's cloud database, and sanctioned the failure to preserve them by preclusion and a fee award. That is a magistrate judge's decision in one district, persuasive rather than binding, and the circuits apply different tests for control. It is still the closest analogue on the books to an ad account.
Retention clocks that run on their own schedule
Nothing in Google Ads, Google Analytics 4, Search Console, Meta Ads Manager or Amazon Ads exposes a control that suspends deletion. There is no litigation-hold switch. The published windows are short relative to the pace of litigation, and the most granular data expires first.
- Individual click records in the Google Ads API, keyed to the click identifier, are available for roughly 90 days, and only one day per query.
- Google Ads change history is viewable in the interface for two years; the API equivalent covers 30 days.
- GA4 user-level and event-level data is retained for 2 or 14 months on the free product, with longer tiers up to 50 months available only on Analytics 360, and expired data "is deleted automatically on a monthly basis" per Google's own documentation.
- Search Console keeps 16 months, rolling.
- Meta's Ads Insights API moved to tiered availability effective 12 January 2026 — 37 months for total aggregates, 13 months for unique-count and hourly breakdowns, 6 months for frequency breakdowns.
- Amazon's Reporting API returns roughly the last 65 days, in requests capped at 31 days.
- Universal Analytics stopped processing in 2023, and access to current and historical UA data ended the week of 1 July 2024.
Read those together and the pattern is clear. In a dispute filed a year after the conduct, the aggregate summary usually survives and the forensic layer usually does not.
Some of what looks like a record is a model output
A growing share of the numbers in these interfaces were never observed. Google's documentation states that modeled conversions "use data that doesn't identify individual users to estimate conversions that Google is unable to observe directly," that "in the 'Conversions' column, Google reports both modeled and observed conversions," and — the sentence that matters most under cross-examination — that the modeling "determines whether a Google ad interaction led to the online conversion. It doesn't determine whether or not a conversion happened." Read precisely, that is a claim about causal assignment, not about existence, and it should not be overstated in either direction. The source is Google's About modeled conversions page.
GA4 does something adjacent: behavioral modeling estimates the behavior of users who declined analytics cookies from the behavior of similar users who accepted, and it appears in standard reports but is excluded from the BigQuery export. Meta told developers in January 2021 that "statistical modeling will be used for certain attribution windows and/or metrics." None of these is an observation of a person doing a thing. An expert who cannot say, for a given figure, whether it was counted or estimated is not yet ready to be deposed on it.
A platform number is a seller's report on its own delivery
The advertising figure in dispute was produced by the party whose performance is being measured, using methods it does not publish, and it is offered as a neutral measurement. Treat it instead as a party's statement about its own product.
There is a documented basis for that framing. Meta published corrections to its own metrics in November and December 2016, disclosing among other things that seven-day and 28-day organic Page reach had been miscalculated as a simple sum — 28-day reach was roughly 55% lower once fixed, an error live since May 2016 — and that Instant Articles average time spent had been over-reported by 7 to 8% since August 2015. Separately, in DZ Reserve v. Meta Platforms (9th Cir. 2024) the Ninth Circuit affirmed certification of a damages class of advertisers alleging that Meta's "Potential Reach" was an estimate of accounts rather than of people, and vacated certification of the injunctive-relief class. The court decided class certification under Rule 23. It did not decide the merits and it made no finding of fraud.
None of that shows any particular number is wrong. All of it means the number needs to be characterized honestly in a report.
Three different records, routinely confused for one
For any paid campaign there are three separate records and conflating them produces bad exhibits. First, the advertiser's own account view: creative, targeting, budget, delivery, spend, reported conversions, change history — complete for that advertiser and available without process. Second, the platform's internal record: bid-level and moderation data the advertiser never sees, obtainable only by subpoena, and resisted. Third, the public archive.
The public archive is where expectations are furthest from reality. Meta's Ad Library launch announcement describes it as covering active ads any Page is running, with ads about social issues, elections and politics also archived for seven years. The spend and reach fields litigators assume are public exist in the Ad Library API only for political and issue ads. And even the seven-year archive is now a rolling window: Meta began removing expired political ads on 24 May 2025, as the first archived year reached the limit. A plan to prove a competitor's ad spend from a public library is a plan that does not work.
Discrepancies are the expected condition, not the anomaly
Two honest extracts of the same period routinely disagree, and an expert who treats every gap as evidence of bad faith will be handed the platform's own documentation on cross. Search Console omits anonymized queries from the table while including them in chart totals — unless a query filter is applied, at which point they drop out of both, so filtered rows cannot be summed to the unfiltered total. Its table caps at 1,000 rows, and its API reference states that it returns top rows rather than all data rows. Aggregating by property versus by page changes the totals.
GA4 collapses long-tail dimension values into an "(other)" row past a table's row limit, applies undisclosed thresholds to demographic and query data, and samples above an event quota that is 10 million events for standard properties. Meta's Insights refresh every 15 minutes and stabilize only after 28 days. A party who says a traffic source "does not appear in the data" may be describing a row limit rather than an absence — and that is a testable proposition, not a rhetorical one.
What these records are good at, and where they stop
Channel and platform records are strong evidence of configuration and delivery: what was set up, by whom, on what date, what was served, what the platform logged, what was charged. That covers a great deal of what is actually contested in agency disputes, scope disputes and compensation disputes, and it is where an analysis is most defensible.
They are weak evidence of receipt, perception and cause. A delivery log shows a server responded. A viewable impression, under the Media Rating Council's 2015 digital guidelines, means at least 50% of the ad's pixels were in the viewable space of an in-focus browser tab for at least one continuous second. Nothing in that definition requires a human to have been present, to have read anything, or to have been persuaded. Every record described on this page is observational; separating what happened from what caused it is a different exercise with different methods.
The single most common failure I see in reports on both sides is a sentence that slides from delivery to persuasion without noticing it did.
What this means in the first weeks of a matter
Because the destruction is scheduled rather than discretionary, the reasonable step available to a party is not "do not delete." It is copy the data out, early, at the finest granularity still offered, and document the extraction. The click-level pull is the most time-critical item in the entire exercise, and it is the one nobody thinks of in the first month.
Two honest habits carry most of the weight. Record, for every extract, the account, the exact query or interface path, the date range, the attribution setting, the reporting identity, the time zone, the currency, and the date and time of extraction — without those, two truthful pulls will not reconcile and the reconciliation will consume a deposition. And write down, contemporaneously, what had already expired when you arrived. "The record could not be found" and "it did not happen" are different statements. A report that conflates them is impeachable on the platform's own documentation, and the gap cuts equally against the party asserting the conduct and the party denying it.
Frequently Asked Questions
Can a litigation hold stop a platform from deleting advertising data?
No. Google Ads, Google Analytics, Search Console, Meta Ads Manager and Amazon Ads have no control that suspends retention, and GA4's documentation states expired data is deleted automatically on a monthly basis. A hold letter binds the parties and their own systems; it has no effect on a third party's schedule. Under FRCP 37(e) the question is whether a party took reasonable steps to preserve, and here the only step within a party's power is exporting the data before it expires. That makes an export protocol, not a preservation instruction, the practical content of a hold over this material.What is the shortest retention window I should worry about first?
Click-level data. Individual click records in the Google Ads API, keyed to the click identifier, are described in Google's Ads API support channel as available for dates back to 90 days before the request, one day per query. There is no documented workaround for older data. Amazon's Reporting API is close behind, returning roughly the last 65 days. Both routinely expire before a complaint is drafted. If a matter involves click quality, invalid traffic, or any question that needs per-click detail rather than campaign totals, that pull is the first thing to schedule.Is data in a platform account within my client's possession, custody or control?
Often, though the answer is not uniform. FRCP 34(a)(1) reaches electronically stored information in the responding party's possession, custody or control, and in Brown v. Tellermate Holdings (S.D. Ohio 2014) the court held that employees with login credentials could retrieve records from a vendor's hosted database and rejected the claim that the company lacked access. That is a magistrate judge's decision in one district and is persuasive rather than binding. The circuits differ on whether control means a legal right or a practical ability, and that split is consequential when the account is held by an agency rather than the client.How do I tell whether a reported conversion was counted or estimated?
Not perfectly, which is itself the finding. Google reports modeled and observed conversions in the same column and does not publish a split for a given metric. In GA4 the clearest available test is comparison against the BigQuery export, because modeled data appears in standard reports but is excluded from that export, from audiences, and from user-level explorations — a figure that cannot be reproduced from the raw export has modeling as a candidate explanation. Beyond that, the honest position is to identify where modeling is present, note the data-quality indicator and its effective date, and quantify the gap rather than the composition.Can the Meta Ad Library show what a competitor spent on advertising?
For ordinary commercial advertising, no. Meta's Ad Library API marks the fields litigators most want — byline, delivery by region, and estimated audience size — as available only for political and issue ads. Meta's own launch announcement describes the library as covering ads a Page is actively running, with political and issue ads additionally archived for seven years. So a commercial ad that stopped running may have no public record at all, and even the seven-year political archive began pruning on 24 May 2025. The advertiser's own account export remains the record that actually contains spend.If two exports of the same period disagree, does that indicate manipulation?
Usually not. Discrepancy is the ordinary condition of these systems. Search Console excludes anonymized queries from its table but includes them in chart totals, caps its table at 1,000 rows, and returns different totals depending on whether data is aggregated by property or by page. GA4 collapses long-tail values into an (other) row, samples above an event quota, and applies undisclosed thresholds. Meta's Insights do not settle until 28 days after reporting. Each of those is documented by the platform, and an expert should reconcile and explain the differences before characterizing any of them as evidence of conduct.What can platform data almost never establish?
That a click or an impression maps to an identified natural person. GA4's reporting identity setting changes user counts from the same underlying data and can be toggled at will with no effect on collection. Google signals data, which drives most cross-device stitching, is capped at 26 months and excluded from the BigQuery export. Amazon Marketing Cloud is architecturally incapable of returning a row representing too few users. A user in these systems is a platform-defined construct, and no ad platform record identifies a person by itself.Published