Evidence and testimony
Abstract tapered column illustration representing Content and Digital Media

EvidenceParty-heldThe record sits in the parties' own systems, and its completeness is itself contested.

Content and Digital Media

Short answer
The publishing party's own systems answer more than any public archive does
The record
CMS revision history, template and container versions, dated captures, server logs
Who holds it
The publishing party; archives and third-party captures are secondary
What it cannot settle
No record shows what a particular reader saw where testing or personalization ran

What a page said on a given day is a harder question than it looks, and the publisher's own systems answer it best

Three different questions get called content disputes

Content and digital media covers material a party publishes to be found, read or watched — articles, landing pages, video, product copy, sponsored placements and the feeds that carry them elsewhere. Disputes in this space almost always reduce to one of three questions, and each has a different record behind it.

Was it produced? That is a deliverables question, answered from the statement of work, the delivery record and the publishing system, and it arises most often between a client and an agency or contractor. Was it published, when, and in what state? That is an evidentiary question about a page at a moment in time, and it is harder than it looks. Did it perform? That is a measurement question, answered from reporting systems whose own policies shape the answer. A digital media expert witness who does not separate the three at intake tends to produce an exhibit that answers a question nobody asked, because the records that settle each one sit in different systems held by different people.

Production and publication live in the CMS, and that record is routinely overlooked

The content management system — the software a site publishes through — is usually the best record in the matter and the least often requested. WordPress keeps post revisions with an author and a timestamp; comparable systems keep equivalents. That record shows what text existed at which point, who saved it under which account, and in what order the changes were made. Where a workflow tool sits in front of the CMS there is a second layer: assignment, draft, review and approval, each dated.

Two adjacent records complete the picture and are almost never collected. The first is the theme or template version history, which governs what the page looked like as distinct from what it said; where a team used a version-controlled workflow that history is complete, and where someone edited the live site in a browser it may not exist at all. The second is the tag manager container version history, which keeps a versioned, timestamped, attributed record of every publish — showing exactly what scripts were live, on which pages, from which date, published by which account. In disputes about what a page did rather than what it said, that container history is the single most probative artifact available, and it is frequently not asked for until it is too late to matter.

What a page said on a given day, and why a screenshot is the weakest way to show it

A screenshot establishes remarkably little standing alone. It does not carry the date, the server's response, the URL as requested, whether the viewer was assigned to a test variant, or whether personalization applied to that session. It records what one browser rendered on one machine under conditions the exhibit does not disclose.

The complication behind that is the experiment layer. If an A/B test was running during the period in dispute, then “what the website said” has no single answer, because different visitors saw different pages by design. Any screenshot-based claim about page content has to be checked against the experiment log for the same period — which variant existed, how traffic was split, when the test started and stopped. That log is unusually good evidence in its own right, because an experiment configuration is contemporaneous and effectively pre-registered: it records what the operator intended before the outcome was known. The counterpart to that strength is that experiment results carry the full set of statistical objections, and an opposing expert will raise all of them — stopping early when the numbers look good, multiple comparisons, insufficient power, sample ratio mismatch, novelty effects, and the distance between a lift on a micro-conversion and any effect on revenue.

Archives and captures: useful, incomplete, and usually assumed to be complete

When the publisher's own records are unavailable, attention turns to third-party archives. The Internet Archive's Wayback Machine is genuinely useful and it is not a complete record. Crawl frequency is irregular, so the gap between two captures may be months. Robots exclusions have historically removed content retroactively. Dynamic and personalized pages are captured poorly or not at all. And embedded resources may be pulled from crawl dates other than the HTML's, which means a reconstructed page can be an assembly of moments rather than a picture of one.

The alternative is contemporaneous capture done for evidentiary purposes. Commercial web-capture services produce captures with hash values and affidavits designed to satisfy FRE 902(13) and 902(14), and the difference between that and a screenshot is the difference between a record with a chain of custody and an image. That is a description of how the tooling is designed rather than an endorsement of any vendor, and the certification it produces establishes authenticity only — not that the captured page was accurate, complete, or representative of what other visitors saw.

Syndication and republication, where the record thins out

Content rarely stays where it was published. It is syndicated to partner sites, pulled through feeds, republished under license, scraped without one, and reproduced inside marketplaces and aggregators. There is no central register of any of that, and the absence is the defining evidentiary feature of the channel.

What exists is held at the origin. The distribution contract or license states what was permitted. The feed configuration and its own logs show what was made available and when. Server logs at the origin — request time, requesting address, user agent, requested path, referrer, response code — are a first-party record on the publisher's own infrastructure, unaffected by any platform's retention schedule or consent gate, and they are the most reliable and least used record in this area. Where republication happened inside a marketplace, the marketplace's own change or contribution log is the primary evidence, which matters in shared-catalog disputes where a third party edits a listing that several sellers depend on. It is worth noting how little the industry's own authorization files help here: the IAB Tech Lab's ads.txt specification applies to websites and treats syndication rules and delegation of third-party authority as future work rather than covered ground. Those files answer who may sell inventory, not who was entitled to republish content.

Sponsored content, and the record a disclosure question actually needs

Where the dispute concerns paid or incentivized content, the FTC's Guides Concerning the Use of Endorsements and Testimonials, 16 C.F.R. pt. 255, are the reference point. The Commission announced revised Guides on 29 June 2023 and published them in the Federal Register on 26 July 2023 at 88 FR 48092, effective on publication. The revision added a principle addressing manipulation of consumer reviews, guidance on incentivized and insider reviews, a definition of “clear and conspicuous” with the observation that platform-provided disclosure tools alone may be insufficient, an expanded definition of endorsements reaching fake reviews, virtual influencers and tags in social posts, and an express statement that intermediaries such as advertising agencies and public relations firms can carry exposure. A precision point that matters in a report: the Guides are administrative interpretations rather than rules, so saying a company “violated the Endorsement Guides” is imprecise — liability runs through Section 5. Separately, the Rule on the Use of Consumer Reviews and Testimonials, 16 C.F.R. pt. 465, took effect on 21 October 2024 and is in force; note that its section 465.3, which would have addressed review hijacking, was not adopted and is codified as reserved.

The marketing expert's contribution here is not the legal conclusion. It is the record: what was published, on what date, in which version; what the disclosure said and where it sat relative to the claim; whether it was visible without interaction on the devices most readers used; and what the publishing and container histories show about when it was added, moved or removed.

Performance evidence, and the reporting policies inside it

Where the question is whether published content performed, the reporting systems answer with their own rules embedded in the numbers, and those rules have to be stated before the numbers mean anything. Search Console keeps data for the last 16 months on a rolling basis. Some queries are omitted from the report to protect user privacy; those anonymized queries are included in chart totals but not in the table, and they drop out of both once a query filter is applied — which is why filtered figures cannot be summed to an unfiltered total, an arithmetic error that gets impeached with the platform's own documentation. The table caps at 1,000 rows, and the API documentation states plainly that it “does not guarantee to return all data rows but rather top ones.” Aggregation by property and aggregation by page produce different totals for the same period.

Analytics adds its own. Where the number of rows exceeds a table's limit, long-tail values are condensed into an “(other)” row — and long-tail values are exactly what a content dispute needs, because the individual article, the individual landing page and the individual referrer are the units in question. A party saying a page does not appear in the data may be describing a row limit rather than an absence. The durable fix in both systems is a bulk export configured to accumulate, and in both it works only prospectively: it preserves everything from the day it is switched on and nothing before it.

What a content analysis does not settle

It does not establish what a particular reader saw. Where an experiment, personalization or a feed algorithm was operating, different visitors received different pages, and the record shows the population of variants rather than one person's experience.

It does not establish authorship in the sense a dispute usually means. A revision history records the account that saved a change at a time, which is a starting point rather than an answer; connecting an account to a person needs access records and corroboration, the same way any account attribution does.

It does not establish that content caused a business outcome. Every record described here is observational, and separating the effect of publishing something from seasonality, competitor activity, pricing, ranking changes and everything else moving at once requires a design that was almost never run. And absence is not proof of absence: a page missing from an archive, a query missing from a report, or a URL condensed into an aggregate row are all facts about the record-keeping system, not about whether the thing existed. Saying which of those is the case, and declining to reason past it, is most of what makes the rest of an opinion credible.

Frequently Asked Questions

How can it be shown what a web page said on a specific past date?

In order of strength: the publisher's own content management system revision history, which records the text, the account that saved it and the timestamp; the template or theme version history, which governs appearance; the tag container version history, which records what scripts were live; a contemporaneous evidentiary capture carrying hash values; and only then a third-party archive such as the Wayback Machine, which has irregular crawl frequency and known gaps. A screenshot with no metadata sits below all of these. Where an A/B test was running, no single answer exists, and the experiment log has to be read alongside whichever record is used.

Is the Wayback Machine reliable evidence of past website content?

It is useful and it is incomplete, and an exhibit should say which. Crawl frequency is irregular, so a capture may be weeks or months from the date in question. Robots exclusions have historically removed content retroactively. Dynamic, logged-in and personalized pages are captured poorly or not at all. Embedded resources may come from different crawl dates than the HTML, so a rendered archive page can combine moments. It is a reasonable secondary source where the publisher's own records are gone, and a weak primary one where they exist and were simply not requested.

What record shows whether a content deliverable was actually produced and published?

The statement of work defines what was owed, and the publishing system shows what appeared. Between them sit the drafts and approvals in whatever workflow tool was used, each dated, and the revision history showing which account saved what and when. Where content was delivered but not published, or published and later removed, the revision and deletion records at the origin usually settle the sequence. Reconstructing from the live site alone is lossy, because the live site shows the current state rather than the history, and that limitation should be stated rather than worked around.

Can an expert establish that published content caused a change in revenue?

Not from the content record alone. Publication records, analytics and search reporting are observational: they show what was published, when, and what was measured afterward. Separating the effect of the content from seasonality, competitor activity, pricing, inventory, ranking changes and concurrent marketing requires an experimental design that in most matters was never run. What the record can support is sequence, magnitude and consistency — what changed, when, and whether the timing is consistent with the claim. Presenting that as correlation with its limits stated is defensible; presenting it as cause generally is not.

What does a disclosure dispute over sponsored content actually turn on?

On the record of what a reader would have encountered. The FTC's Endorsement Guides, revised in 2023, added a definition of clear and conspicuous and observed that platform-provided disclosure tools alone may be insufficient. Translating that into evidence means establishing what was published and when, what the disclosure said, where it sat relative to the claim, whether it was visible without interaction on the devices most readers used, and what the publishing and container histories show about when it was added or changed. The Guides themselves are administrative interpretations rather than rules, and a report should describe them that way.

Why do content performance reports from two systems disagree?

Because each applies rules that are documented but rarely read. Search Console holds 16 months on a rolling basis, omits anonymized queries from the table while including them in chart totals, drops them from both when a query filter is applied, and caps the table at 1,000 rows; its API does not guarantee to return every row. Aggregating by property rather than by page changes the totals. Analytics condenses long-tail values into an aggregate row past a limit. Discrepancies between two honest extracts are the expected condition rather than an anomaly requiring an explanation of bad faith.

What should be preserved first in a content or digital media dispute?

The content management system's full revision history for the pages at issue, exported rather than screenshotted. The theme or template version history and the tag container version history, both of which are small, complete and easily lost in a redesign. Origin server logs for the relevant period, which are first-party and unaffected by any platform's retention schedule. Experiment and personalization configurations covering the same dates. Then the reporting layer, which expires: Search Console holds 16 months, and a bulk export configured now preserves everything from now forward and nothing before it.
Keep reading

Read the guides

An entry names the record that exists for one channel or one claim. A guide covers what is done with it, and how long there is before a retention window closes.

Top