Prerequisites
01
Deleted is not gone: archived copies survive long after the original disappears
Web archives, search-engine caches and third-party preservation services capture content continuously and independently of the original publisher. When a subject deletes a page, post or account, investigators with the right recovery workflow can still retrieve what was there, and prove it.
Archive recovery is the systematic retrieval of web content that no longer exists at its original URL, using copies held by independent archiving services. The technique applies across deleted websites, removed social media posts, altered news articles and withdrawn documents. Recovery is possible because archiving services operate on automated crawl schedules that are independent of the original publisher's decisions.
In the field
On 17 July 2014, shortly before news broke that Malaysia Airlines Flight MH17 had been shot down, Donetsk separatist commander Igor Girkin posted a claim on VKontakte that his forces had downed a Ukrainian military transport plane. The Wayback Machine, which had listed Girkin's VK page for regular archiving two weeks earlier at the request of a Hoover Institution curator, captured the post minutes after it went up. Girkin's page was edited to remove the claim within hours, but the Internet Archive capture survived and was independently corroborated by multiple news organisations.
- Wayback Machine snapshot retrieval. A scheduled crawl captured Girkin's VK page at 15:22 GMT, roughly 30 minutes after the post was made, preserving the original text and video links.
- Timestamp comparison. The capture timestamp predated the page edit by roughly two hours, establishing that the claim was live and publicly visible before it was removed.
- Cross-source corroboration. Multiple independent outlets obtained or referenced the same captured text, corroborating the Wayback Machine record independent of any single source.
Internet Archive Wayback Machine · Christian Science Monitor · MH17 investigation · 17 July 2014
Learning outcomes
By the end of this tutorial you will be able to:
Retrieve archived versions of deleted web pages using Wayback Machine and alternative archive services
Recover ephemeral content including deleted social media posts using platform-specific cache and archive tools
Establish the date of a page version using archive timestamps and compare versions across time
Document archived content to evidentiary standard with chain-of-custody records
Identify the limits of archive coverage and determine when a gap in the record is evidentially significant
The rest of this tutorial is for Signal subscribers.
What remains: the decision framework, the tool configuration, the failure modes, and the evidentiary standard required to use the finding defensibly. Signal is €90 a year, or €9 a month. Students, €49 a year.
Join SignalA Signal subscription gives you:
- Full OSINT Reference Card library, 21 domains
- Methods, every tradecraft tutorial in full
- AI in OSINT, every prompt and field report
- Shadow Analysis, every forensic dossier


