Evidence and testimony
The Evidence Behind a Marketing Claim

What Platform Data Does and Does Not Record

The gap between what an advertising platform shows you and what it actually recorded is where most reports go wrong

The record is a set of events the platform chose to keep, in a shape it chose

Every figure in an advertising or analytics interface is the output of a definition the platform wrote and can change. That is not a criticism; it is the operating reality, and it has to be stated in a report rather than assumed away.

Google Analytics 4 is the clearest illustration. It is event-based — Google's documentation says "any interaction can be captured as an event" and that "in Google Analytics 4 properties, every 'hit' is an event; there is no distinction between hit types." Universal Analytics, its predecessor, was hit- and session-based, and derived sessions differently. Google's own guidance was to rethink data collection in the new model rather than port the old structure across. The practical consequence for a dispute spanning the changeover is that a session count before and a session count after are not the same quantity, and any year-over-year exhibit crossing that boundary is measuring a model change alongside whatever else it claims to measure.

Worse, for conduct measured in Universal Analytics before mid-2023, the contemporaneous record is gone. Access to current and historical UA data ended the week of 1 July 2024. Whatever survives is what a party exported first — spreadsheets, PDFs, screenshots, a warehouse copy — and those are secondary artifacts that have to be authenticated as such.

Counted, estimated, and reported in the same column

The single most consequential thing an attorney can learn about this data is that observation and estimation are mixed together without a visible seam.

Google states that modeled conversions "use data that doesn't identify individual users to estimate conversions that Google is unable to observe directly," and that "in the 'Conversions' column, Google reports both modeled and observed conversions." It applies modeling where third-party cookies are unavailable, where consent is absent, where device tracking is restricted, and across cross-device journeys. Google's documentation also draws a narrower line than critics usually allow, stating that the modeling "determines whether a Google ad interaction led to the online conversion. It doesn't determine whether or not a conversion happened."

GA4 runs a parallel mechanism. Behavioral modeling for consent mode "uses machine learning to model the behavior of users who decline analytics cookies based on the behavior of similar users who accept analytics cookies," and it only engages above documented eligibility thresholds — at least 1,000 events per day with analytics storage denied for at least 7 days, and at least 1,000 daily consenting users for 7 of the previous 28 days. Below those thresholds the gap is simply missing rather than filled. Meta told developers in January 2021 that "statistical modeling will be used for certain attribution windows and/or metrics."

The one clean test for whether modeling is in a number

Google does not publish a modeled-versus-observed split for a given metric in the GA4 interface. What it publishes is where modeled data goes. Modeled data is included in standard reports, affecting users and sessions, and is explicitly excluded from audiences, User explorer, cohorts, segments with sequences, retention reports, and the BigQuery export.

That exclusion is the test. A figure that appears in a standard report but cannot be reproduced from the property's raw event export has modeling as a candidate explanation, and the size of the difference is measurable even though its composition is not. The interface signals the condition with a data-quality indicator carrying wording such as "including estimated user data" together with a modeling effective date — a per-card banner, not a per-number breakout.

An expert can honestly do three things with this: identify where modeling is present, state the effective date, and quantify the gap. An expert cannot validate the model, reproduce it, or say for any individual number how much of it was counted. Whether platform-modeled figures should be treated as ordinary business records going to weight, or as outputs of an undisclosed proprietary model produced by the party whose performance is at issue, is genuinely contested and unresolved by any published decision I have located.

What suppression removes before you ever see the report

Several systems withhold rows for privacy reasons, and the withholding is invisible unless you know to look for it.

  • GA4 data thresholds. Applied to reports containing demographic data, audiences built on demographics, and search-query information where the user count is insufficient. Google states the thresholds are "system defined. You can't adjust them," and does not publish the value. Documented mitigations are widening the date range or exporting to BigQuery — though Google Signals data is excluded from that export.
  • Search Console anonymized queries. Per Google's dimensions documentation, "some queries are omitted from the report to protect user privacy." They are included in chart totals but omitted from the table, and drop out entirely once a query filter is applied.
  • Google Ads search terms. The report lists only "search terms that a significant number of people have used," and terms without enough query activity are omitted. Coverage was restricted in September 2020, partially restored in September 2021 for queries from 1 February 2021 forward, and pre-September-2020 sub-threshold data was removed as of 1 February 2022.
  • Row limits. GA4 condenses less common dimension values into an "(other)" row past a table's limit; Search Console's table caps at 1,000 rows and its API allows up to 25,000 while its reference states that it returns top rows rather than all data rows.

The search terms sequence deserves particular care in trademark and competitor-bidding matters. A time series crossing those dates measures Google's reporting policy at least as much as it measures anyone's behavior.

A user is a setting, not a person

GA4 offers three reporting identities — blended, which resolves by user ID, then device ID, then modeling; observed, which uses user ID then device ID; and device based, which "uses only the device ID and ignores all other IDs that are collected." Google states that the choice "does not affect data collection or processing" and that a party can switch between them at any time without permanent effect on the data.

Read that carefully. The same property can produce different user counts for the same date range depending on a setting one party controls, and switching it leaves no mark on the underlying data. A screenshot of a user count is therefore not self-authenticating as to which identity space produced it, and any exhibit resting on unique users needs the setting recorded alongside the number.

Cross-device resolution has its own limit: Google signals data, which drives most of it, is capped at 26 months regardless of the retention setting, respects a shorter setting if one is configured, and is excluded from the BigQuery export. On the commerce side, Amazon Marketing Cloud is a clean room that "only accepts pseudonymized information," returns only aggregated outputs, and enforces an aggregation threshold that suppresses rows representing too few users. It is useful for measurement questions and architecturally incapable of answering identification questions.

What the platform records reliably, and it is more than people expect

Not everything here is fog. Configuration and change are recorded well, and those records are frequently the most probative material in a dispute about competence or scope.

Google Ads change history covers the last two years in the interface and logs changes to ads, assets, audiences, budgets, bid adjustments, conversions, feeds, targeting, keywords, and campaign and ad group status, including changes made through automated rules, the API and Google Ads Editor, attributed by user. It does not track password changes. The API equivalent reaches back only 30 days, which is why the interface capture matters.

GA4 records destructive acts. A data-deletion request replaces collected event-parameter text with "(data deleted)" while the events still count toward metrics; it requires an Editor role, has a seven-day cancellable grace period, processes in 7 to 63 days, cannot touch data less than 12 days old, and is irreversible once complete. Property and account deletion runs through a 35-day trash window, after which Analytics "permanently deletes the entity, and records the deletion in Change History, with the user identified as the Analytics System." Tag manager version history — who published which tag, on which pages, from what date — is equally good and is routinely not collected.

The invalid-traffic columns, and what they are counting

Google defines invalid clicks as "clicks on ads that aren't the result of genuine user interest, including intentionally fraudulent traffic and accidental or duplicate clicks," and states that advertisers are not charged for them. Advertisers can add an "Invalid clicks" column covering traffic received over roughly the last 60 days, and Report Editor offers an Invalid Activity Credit Report showing credited clicks, credited interactions, credited amount and adjusted performance by campaign and network.

Both are records of what the platform's own systems rejected. Neither carries per-click reason codes, actor attribution, or any denominator for what the filters missed. And the credit bucket is mixed: Google states the credits reflect invalid traffic or interactions later determined to be on inventory violating an AdSense program policy, so a credited click is not necessarily a click Google classified as fraudulent. Google also documents that a click may be removed as invalid while the conversion arising from it is not, which "can occasionally result in conversions being higher than clicks" — an anomaly with a published innocent explanation.

The remedy is documented too, and it constrains a damages theory: "clicks determined to be invalid will result in adjustments or credits, not a refund."

Moving tagging server-side moves the custodian

Server-side tagging routes hits through infrastructure the advertiser provisions and serves from its own domain, using the same tag, trigger and variable model as a browser container. Meta's Conversions API does the analogous thing, connecting an advertiser's server, platform or CRM to Meta's endpoint.

This changes the evidentiary picture in three ways at once. It creates a party-held copy of the event stream where none existed, which is the best available evidence in a measurement dispute. It moves the discovery target, so the answer to "prove the conversion fired" becomes server logs rather than a platform report — and Google's own server-side tagging material notes that the reference Cloud Run deployment writes standard-output and request logs unless request logging was disabled, which makes "there are no server logs" a discoverable assertion rather than an architectural fact. How long those logs survive is a question for that advertiser's cloud logging configuration, not for Google's tagging documentation.

And it inserts a party-controlled transformation step ahead of the platform. A container can add, drop, hash or synthesize parameters before forwarding, so a discrepancy between advertiser and platform data can originate in the advertiser's own configuration. Server-side data is better evidence and a new place for error, and it deserves the same chain-of-custody discipline as any other party-held log.

What the platform record does not settle

Four things, stated plainly, because an opposing expert will raise every one.

Absence in the data is not absence of the event. Every system described here destroys on a schedule. By the time counsel is engaged the granular record has often been deleted by the platform rather than by a party, and that unavailability is symmetrical between the parties.

A platform figure is not an independent measurement. It is the seller's report on its own delivery, produced by methods it does not publish. Meta's own newsroom corrections in November and December 2016 disclosed that several of its metrics had been materially misstated for months.

No record here identifies a natural person, and the same underlying activity yields materially different user counts under different settings.

Nothing in this material establishes cause. These are observational records of delivery and configuration. A campaign record can be complete, accurate and fully authenticated and still say nothing about whether the campaign moved revenue — and a damages model built on modeled conversions, thresholded rows, sampled explorations and a shifting attribution window inherits every one of those limits rather than escaping them.

Frequently Asked Questions

Does a GA4 export show everything the property collected?

No, and which parts are missing depends on the export. The BigQuery export is the raw event data Google Analytics receives from the client, which makes it the closest thing to a primary record, but it excludes modeled data and Google Signals data, it has a daily volume limit of one million events on standard properties, and it does not backfill — it starts when it is configured. Standard reports include modeled data and are not affected by the retention setting, while explorations and funnel reports are. Any comparison between the two needs those differences stated.

Why do Search Console totals fail to add up?

By design, and Google documents it. Anonymized queries are omitted from the report table but included in chart totals, and applying a query filter drops them from both — so filtered rows cannot be summed to an unfiltered total. The table caps at 1,000 rows, Search Console stores only the most important rows rather than everything it collects, and the API returns top rows rather than a complete set. Aggregating by property and aggregating by page produce different figures for the same period. An expert who performs that arithmetic in a report will be impeached with Google's own page.

Can an expert tell how much of a conversion count was modeled?

Not as a percentage of any specific number. Google reports modeled and observed conversions in the same column and publishes no split. What can be done is bounded and defensible: identify that modeling is present from the data-quality indicator and its effective date, note the conditions Google says trigger modeling, and measure the difference between a standard report figure and the same figure reconstructed from the raw event export, since modeled data is excluded from that export. That difference is a gap with a candidate explanation, not an allocation, and it should be described that way.

Does the Invalid clicks column measure fraud committed against the advertiser?

It measures what the platform's own filters rejected. It has no reason codes, no per-click detail, no actor attribution, and no denominator for what was missed, so a low figure is equally consistent with clean traffic and with undetected sophisticated activity. The separate Invalid Activity Credit Report is the more useful export because it shows credited clicks and credited amounts by campaign and network, but Google states that bucket also includes interactions on inventory that violated an AdSense program policy. A credited click is therefore not necessarily a click the platform called fraudulent.

What record shows who changed a campaign and when?

Google Ads change history, which covers two years in the interface and logs changes to ads, assets, audiences, budgets, bid adjustments, conversions, feeds, targeting, keywords and status, attributed by user, including changes made through automated rules, the API and Google Ads Editor. It does not track password changes, and the API version reaches back only 30 days, so the interface capture is the one worth taking. On the site side, tag manager version history records which container version was published, by which account, on what date — an equally strong record that is rarely requested in discovery.

Is server-side tracking data more reliable because it is first-party?

It is better positioned, not automatically more reliable. Routing events through infrastructure the advertiser controls creates an independent copy of the event stream and puts it plainly within the advertiser's control for discovery purposes. It also inserts a step where parameters can be added, transformed, hashed, dropped or synthesized before the platform sees them, which means a discrepancy between advertiser and platform figures can originate in the advertiser's own container configuration. Treat a server-side event stream as party-held evidence requiring the same chain-of-custody discipline as any other log.
Keep reading

The entries behind this guide

Every channel, record type and claim named here has its own entry: where the record lives, who holds it, and what it cannot settle.

Top