Marketing · 13 min read · Updated 7 Aug 2026

How to See a Competitor's Old Website: Web Archives, Ranked

Every other source in this cluster tells you what a competitor says now. Web archives tell you what they said before, and they are the only source that produces evidence a third party can check for themselves: a snapshot sits at a public, dated address on somebody else's servers, which is something your own screenshot will never be. The limit matters just as much. An archive is a sample of a website rather than a record of it, so a snapshot proves what a page said on the day it was captured and proves nothing at all about the days in between.

What web archives contain, and what they do not

A web archive stores what a crawler received when it asked for an address on a particular day: the HTML, and whatever assets it managed to fetch alongside it. The largest of them has been doing this since 1996 and passed one trillion web pages in October 2025, marking the milestone publicly on 22 October that year. For competitive work that turns a competitor’s website from a description of the present into a dated series, which is the only way to see a decision rather than a result.

What an archive does not contain is a record. It holds captures, taken irregularly, weighted toward addresses that were well linked, missing anything behind a login or a form, and frequently unable to reproduce a page that assembles itself in the browser. Every useful thing on this page follows from holding those two facts at once: the evidence is unusually strong, and the coverage is unusually uneven.

Eleven places a competitor's past is stored, what each holds and how reliably you can reach it
SourceWhat it gives youCostHow currentReliability
The Wayback Machine calendarEvery capture the archive holds for one address, by year and day, from 1996 onward and free without an accountFreeCrawl-dependent, from hourly to never High
The Wayback Machine changes comparisonA side-by-side difference between any two captures of the same address, with added and removed text markedFreeAs current as the two captures High
Save Page NowA capture you create yourself, at a permanent public address with today's timestamp on itFreeImmediate High
archive.todayAn on-request capture rendered by a full browser, which preserves the visual page where script-heavy sites defeat other archivesFreeOnly what somebody asked for Medium
Search engine cachesNothing any more. The largest search engine retired its cache links and the cache operator in 2024 and now points at the archive insteadDiscontinuedDiscontinued Low
The site's own surviving URLs and redirectsOld pages that still resolve because nobody removed them, and redirects that map a retired page onto its replacementFreeLive High
Their changelog, release notes and press archiveA dated history the company maintains itself, usually more complete and more precise than any crawl of itFreeLive High
App store version historyEvery release of a mobile product with its date and notes, which is a product history no web crawler assemblesFreePer release High
National library web archivesDeep, curated coverage of one country's domains under legal deposit, including sites the global archives crawled thinlyFree, some access on-site onlyAnnual to quarterly High
Curated institutional collectionsThemed collections built by libraries and archives, searchable by the full text of the captured pages rather than only by addressFree to browseSet by the collecting institution Medium
Their public repository historyFor sites published in the open, the actual edit history with authors and dates, which is better than any reconstructionFreeLive High

How to see a competitor's old website, step by step

  1. 1Start from the exact address, not the home page. Archives index addresses rather than sites, so the coverage of a pricing page and of the home page it links from can be completely different. Paste the deep link to the page whose history you want. If it returns nothing, try the address it used to have, since a redesign usually moved it.
  2. 2Read the calendar before reading any snapshot. The calendar shows how densely the address was captured, and that shape is itself information. Dense years mean the page was popular or heavily linked, and a gap means the archive simply was not there, not that the page went unchanged. Pick your comparison dates from where the captures actually are.
  3. 3Compare two captures rather than reading one. A single old page is a curiosity. The value is the difference, and the archive will compute it for you: choose two dates and it marks what was added and what was removed. Start with roughly a year apart, then narrow around whichever interval shows the change you care about.
  4. 4Check whether the page actually replayed. Archived pages frequently render incompletely, and an empty pricing table is far more often a broken replay than a removed price. View the source of the snapshot as well as the rendered version, because the text is usually still in the file even when the layout collapses.
  5. 5Try a second archive when the first one fails. The archives crawl differently and hold different things. A request-based archive that renders with a full browser often has the page a crawler-based one could not capture, a national library may hold a local site the global archives skipped, and a curated institutional collection can be searched by what a page said rather than by its address.
  6. 6Capture today what you will want to cite next year. The one part of this source you control is the future. Any page you may later need to prove said something can be saved on request in seconds, which puts a public dated copy at a permanent address. Do this for a competitor's pricing, claims and positioning pages the moment you first read them.
  7. 7Record the snapshot address, not a screenshot of it. The archived address contains its own timestamp and is the whole reason this source is different from every other one here. Paste that address into whatever you write, so anybody who doubts the finding can open the same page rather than take your word for what a picture shows.

Why an archived page is evidence and a screenshot is not

Almost everything a competitive team collects has the same weakness: it is an assertion by the person who collected it. A screenshot in a slide is a picture you made. A note saying their pricing changed in March is your memory. Both are fine internally and both collapse the moment somebody with an interest in the answer decides to push back, which is exactly when this work matters most.

Archives are different in one structural way. A capture lives at a public address that contains its own timestamp, hosted by an organisation with no stake in your argument, and anybody who doubts you can open it themselves in five seconds. That is provenance rather than assertion, and nothing else in this cluster produces it. Archived pages have been accepted as evidence in courts and by regulators in several jurisdictions, usually alongside testimony about how the archive works, which is a far higher bar than a competitive brief needs to clear.

The practical rule that follows

Cite the archived address, never a picture of it. Paste the full snapshot link into whatever you write, so the claim carries its own proof. The moment you convert it into an image you have thrown away the only property that made this source different, and you are back to asking colleagues to take your word for it.

This matters most in three places: an advertising or comparison claim a rival has since quietly revised, a capability they used to advertise and no longer do, and a price that moved. In all three the argument is about what was said, not about what is true, and a dated third-party copy settles it in a way nothing else will.

Why web archives are a sample of a website, not a record of it

Capture frequency broadly follows how prominent and how heavily linked an address is. A large company’s home page can be captured many times a day; a mid-sized competitor’s feature page might be captured twice a year; a page nobody ever linked to may have no captures at all. Nothing about that is a judgement on the page’s importance to you.

The consequence is worth writing on the wall, because it is the error that turns this source from useful into misleading. An archive can show you that a page said something on a given date. It can never show you that the page did not say something else in between two captures. A claim added in April and removed in June, with captures in March and July, leaves no trace whatsoever, and the archive will look complete while hiding it.

Two habits follow. First, always report the interval rather than the date: the pricing page changed between 12 March and 8 July, not on 8 July. Second, read the calendar as data in its own right. A competitor whose site is captured densely has a lot of inbound links and a lot of attention, and one whose captures are sparse does not, which is a rough but free reading of web presence that you got on the way to something else.

What breaks when an archived page replays

Archived pages frequently render badly, and people conclude the content is gone. It usually is not. A capture stores the files the crawler could fetch, and a modern page needs considerably more than that, so the failure is almost always in the reassembly rather than in the archive.

Five ways an archived page fails to replay, what actually happened, and how to recover the content
What you seeWhat actually happenedWhat to do
An empty pricing tableThe prices were assembled in the browser after load, so they were never in the captured responseRead the snapshot's source rather than its rendered form. The figures are often sitting in a data block inside the file
Unstyled text on a white pageThe stylesheet was served from a domain the crawler did not captureNothing is lost. The content is intact and reading it as plain text is frequently faster anyway
A consent banner covering the pageThe capture happened before the banner was dismissed, and the crawler had no way to dismiss itTry the capture either side of it, or read the source, where the underlying page is usually complete
Missing images and logosAssets sat on a separate domain or a CDN the crawler did not followCheck whether the alt text survived, which is often enough to identify a logo wall or a screenshot
A redirect to the current siteThe capture stored a redirect rather than a page, usually taken after the address was retiredStep back to an earlier capture, or try the address the page had before the redesign

One more failure has no workaround and needs stating: a crawler that arrived during an experiment was served one arm of it, and nothing in the capture tells you that. So a single snapshot of a heavily tested page, a home page or a pricing page above all, is one variant that existed rather than the page everybody saw. Where the conclusion matters, check two or three captures from the same period before treating any of them as canonical.

Every web archive worth using, and how to work it

1. The Wayback Machine calendar

The default and the deepest. Paste the full address of the page you want, not the domain, and read the calendar before opening anything. The density of the captures tells you how precisely you can date any change, which is the ceiling on every conclusion you will draw. If the address returns nothing, try the address the page had before their last redesign, since a redesign usually moved it rather than deleting it.

2. The Wayback Machine changes comparison

The most under-used feature in this whole area. From the calendar view, open the changes option, select two captures, and it renders them side by side with the added and removed text marked. It converts an hour of squinting at two tabs into a minute, and it is particularly good on pages where the change is a single clause: a dropped guarantee, a removed integration, a softened claim. Start about a year apart to locate the interval, then narrow.

3. Save Page Now

The only part of this source you control, and the reason to bookmark it. Paste an address and the archive creates a capture immediately, at a permanent public location with today’s timestamp. Make it a habit for competitor pricing pages, homepage claims and any comparison page that names you, because those are precisely the pages that get quietly revised and the ones you will later wish you could prove.

4. archive.today

A request-based archive that renders each page with a full browser and stores the visual result. That makes it markedly better than a crawler on script-heavy pages and on anything where what you need is how the page looked rather than what was served. Its coverage is entirely dependent on somebody having asked for that page, so it is thin as history and strong as a second attempt when the first archive returns a broken shell.

5. Search engine caches

Gone, and worth knowing so you stop looking. The largest search engine removed its cached-page links in early 2024 and retired the cache operator shortly afterwards, with its search liaison explaining that the feature had been built for an era when pages routinely failed to load. In September 2024 it added a link to the Internet Archive inside its about-this-page panel, which is now the nearest equivalent workflow. The practical effect is that one archive carries far more of this job than it did two years ago, which is the argument for having a second.

6. The site’s own surviving URLs and redirects

Before reaching for an archive, try the old address on the live site. Companies routinely leave retired pages up because nobody removed them, and where they did tidy up, the redirect tells you which page replaced which. That mapping is genuinely useful on its own: a set of feature pages all redirecting to one platform page is a consolidation, and it dates a repositioning without any archive involved.

7. Their changelog, release notes and press archive

The company’s own history, maintained by them, usually more complete and always more precise than a crawl of it. A changelog gives you dated capability changes without any inference, and a press archive gives you the announcements in their own framing. Read these first for anything product-related, and use the archive for the pages they would rather you did not compare.

8. App store version history

For any competitor with a mobile product, the store holds a dated list of every release with its notes. That is a product history nobody had to crawl, it goes back years, and the cadence alone tells you how much engineering attention the mobile product actually gets. Release notes are also written for users rather than for search engines, which makes them unusually direct.

9. National library web archives

Several countries archive their own domain under legal deposit rules, and the coverage of local sites is frequently deeper than the global archives manage. This is the right place to look for a competitor whose main site is on a country domain, or for a regional business the global crawlers only visited occasionally. Access rules vary, and some collections are readable only from a library reading room, which is a real constraint rather than a formality.

10. Curated institutional collections

Libraries, universities and public bodies build themed web collections, and the largest of these programmes exposes its holdings publicly across well over a thousand collecting organisations. The reason to care is the search: these collections can be queried by the full text of the captured pages rather than only by address, which is the one thing the ordinary archive lookup cannot do. If you need to find who was saying a particular phrase in 2018 rather than what one address said, this is where to start.

11. Their public repository history

Where a company publishes its website or documentation in the open, the actual edit history is available with authors, dates and the exact wording before and after. That is better than any reconstruction, because it records the change rather than sampling the result. Common for developer-facing companies and for documentation sites generally, so check it first when the competitor sells to engineers.

What web archives cost, and the limits worth knowing

  • Reading and saving are both free, with no account. The largest archive is run by a non-profit library funded by donations rather than by subscriptions, which is why nothing here has a paywall and also why nobody owes you an uptime commitment on the day you need it.
  • Robots files are no longer treated as an archiving instruction. The archive said publicly in 2017 that rules written for search engines work badly for archives, having already stopped honouring them on government and military sites some months before. Removal is handled as a request to the archive instead, so a site currently blocking crawlers may still have years of history in it.
  • Nothing behind a login or a form is in there. Product interfaces, gated content, customer portals and anything requiring a search query were never reachable by a crawler. This is the single largest gap and it does not close.
  • Coverage is by address, not by site. A domain being well captured tells you nothing about the specific page you want. Always start from the deep link, and check the calendar for that address rather than assuming the site’s general coverage applies.
  • Work at reading pace, one page at a time. These are free public services run on donations. Open the captures you need, save the ones you care about, and treat the archive the way you would treat a library rather than a database.

What an archived page is commonly misread as proving

Six confident conclusions drawn from a snapshot, and what the snapshot actually supports
What people read it asWhat it actually is
They never made that claim, the archive has nothingThere is no capture containing it. Absence of a capture is not absence of a page, and thin coverage is the normal case rather than the exception
They changed the price on 8 JulyIt changed between two captures. The date you hold is the date a crawler visited, and the event sits somewhere in the gap before it
This is what the page looked likeThis is what a crawler was served: no consent state, no personalisation, no session, and whichever arm of a running experiment it happened to get
The page is gone, so the product was droppedPages are moved, merged and renamed in every redesign. Check the redirect on the live site before reading a removal as a decision
The archive is a complete record of the webIt is a sample weighted toward well-linked addresses, and it holds nothing that sat behind a login, a form or a search box
I took a screenshot, which is the same thingA screenshot is an artefact you made and asked colleagues to trust. The archived address is a dated third-party copy anybody can open

Which competitor questions web archives can answer

Competitor questions a page history answers well, answers weakly, and hands to another source
The questionHow well web archives answer itCovered in full
Who were their customers a year agoVery well. A logo wall is well linked, so it is captured often and the differences are readablecompetitor customers
How has their positioning changedBetter than any other source, because the wording is preserved rather than rememberedcompetitor positioning
What did their product do beforePartly. Feature pages survive; the product behind the login never doescompetitor roadmap
What was the site built onWell. Captures carry the generator tag and asset paths that were live that daywebsite technology
Who used to run which functionWeakly. Team pages are captured unevenly and are frequently out of date when they arecompetitor leadership team
Which partnerships did they announce and dropWell. Partner directories and press pages are both well linked and heavily capturedcompetitor partnerships

The material was published to the public web deliberately, and an archive is a library holding a copy of it. Reading that copy is unremarkable. The care needed here is almost entirely about how you present what you found, because an old page is a true statement about the past and a false one about the present.

  • Never present a withdrawn claim as a current one. A competitor who advertised something in 2023 and stopped is not making that claim today. Repeating it to a market as though they were is misleading comparative advertising, and the fact that you sourced it honestly does not help you.
  • Date every quotation at the point of use. Put the capture date in the sentence, not in a footnote. A battlecard line reading “as of March 2024 their pricing page stated” is usable indefinitely; the same line without a date becomes wrong the moment they change it and nobody notices.
  • Copyright survives archiving. An archived page carries the same rights as the original. Short extracts with attribution are ordinary practice; reproducing a competitor’s page wholesale into your own materials is not, and no research purpose changes that.
  • Old pages can contain personal data. Team pages, testimonials and author bylines name people who may have left, and privacy rules attach to that information regardless of where you found it. Keep what has a business purpose in a professional context, and drop the rest.
  • Do not go looking for what was never meant to be public. Archives occasionally hold a page that was published by mistake and removed quickly. Deliberately hunting for material a company took down for confidentiality reasons is a different activity from researching their public positioning, and it is the version that gets a company into trouble.

What web archives will not show you, and the best proxy

  • Anything behind a login. The product itself, the customer portal and gated content were never reachable by a crawler. Proxy: their documentation and release notes, which describe the product publicly and are dated.
  • What happened between two captures. A change that was reversed leaves no trace at all. Proxy: save the page yourself on a fixed rhythm, which converts a sample into something closer to a series.
  • Why anything changed. The archive records wording, never reasoning. Proxy: what else moved in the same fortnight, since a pricing change alongside a new tier and three sales adverts is a strategy rather than an edit.
  • Which version real visitors were shown. Experiments, personalisation and geography all mean the crawler saw one page among several. Proxy: several captures from the same period, and your own visits from different regions.
  • Small competitors in any depth. Capture density follows links, so the private rival you most want to track is often the one with four captures a year. Proxy: their own changelog and press archive, plus saving their key pages yourself from today onward.

How to keep archive research current

The rhythm here is backwards compared with every other source in this cluster. There is no point checking the archive monthly, because it will not have added much and nothing you already read will have changed. What is worth doing monthly is the opposite: save the five or six pages per competitor that decide arguments, their pricing page, their comparison page, their homepage and their customer list, so that a year from now you own a dense series rather than whatever a crawler happened to take.

Do the actual reading quarterly, and only against a question. Comparing two captures a quarter apart with no hypothesis produces a list of cosmetic edits. Comparing them because you want to know whether their entry tier moved, or whether a named integration disappeared, produces an answer in ten minutes. Record findings as intervals with both timestamps, and file them where the rest of the price history lives, which for most teams is a pricing teardown rather than a separate archive document.

Why web archives arrive late, and how to automate the watch

The structural problem with this source is that it is somebody else’s crawler on somebody else’s schedule. A competitor changes their entry price on a Tuesday; the next capture lands in nine weeks; you notice in the following quarter’s review, by which point three deals have been lost to a number nobody on your side knew about. The archive did its job perfectly and still told you months late, because it was never designed to alert anybody. It is a library, and libraries are for afterwards.

Watching a handful of pages across several competitors and knowing the day they change is a different function entirely, and it is what competitive intelligence software exists to do. Flares keeps its own dated record of what moved on the surfaces a competitor controls, so the history belongs to you and arrives as it happens rather than whenever a crawler next passes. What that does not replace is the archive itself. No monitoring product can show you a page from 2019, and the third-party timestamp that makes an archived capture citable to somebody who does not trust you is something only a neutral archive can provide.

See changes web archives capture months later

Flares records a competitor page the day it changes, rather than whenever the next crawl happens to arrive.

Discover Flares

14-day free trial · 30-second setup

Marketing sources FAQ

How do I see what a competitor's website used to look like?

Open the Wayback Machine, paste the full address of the specific page you want rather than the home page, and pick a date from the calendar it returns. For a page that never got captured, or one that renders as an empty shell, try a request-based archive such as archive.today, which renders with a full browser and often preserves what a crawler could not. Work from deep links throughout: the coverage of a pricing page, a careers page and a home page on the same domain is frequently very different.

Is the Wayback Machine free to use?

Yes, entirely, with no account and no payment, for reading and for saving pages. It is run by the Internet Archive, a non-profit library funded by donations and by its digitisation services rather than by subscriptions. That is worth knowing for two reasons: there is no commercial incentive shaping what gets kept, and the service occasionally slows or goes down without anybody owing you an uptime commitment. Plan around the second point by saving anything critical rather than assuming it will be reachable on the day you need it.

How far back does the Wayback Machine go?

To 1996, and by October 2025 the Internet Archive had passed one trillion web pages archived, a milestone it marked publicly on 22 October that year. Depth is not the same as reach, though. For any specific address the practical answer is whatever the calendar shows, and for a small company that answer is often a handful of captures a year starting from whenever somebody first linked to them. The trillion is the corpus; your competitor's page history is a much smaller and more uneven thing.

Why is a page missing from the Wayback Machine?

Usually because nothing ever pointed a crawler at it. Archives reach pages through links, so deep pages on low-traffic sites, anything behind a form or a login, and anything published and removed inside a single crawl interval can all be absent. Other causes are a page that was never linked internally, a site that blocked the archive's crawler at the time, and a removal request from the site owner. Absence is not evidence that a page did not exist, which is a distinction worth stating explicitly whenever you report a gap.

Can I see a competitor's old pricing page?

Often, and it is the highest-value use of this source. Pricing pages get linked from all over the web, so they are captured more densely than most of a site, and comparing two captures a year apart gives you the price history, the packaging changes and the moment a tier appeared or quietly disappeared. Two cautions. Prices rendered by script frequently do not survive a crawl, so an empty table usually means a broken replay rather than a removed price. And the archived figure is list price. The rest of the sources for a real number are under competitor pricing.

Does the Wayback Machine save every page of a website?

No, and treating it as though it does is the main way this source misleads people. It samples. Capture frequency broadly follows how prominent and how linked an address is, so a large company's home page may be captured many times a day while a mid-sized competitor's feature page is captured twice a year. The consequence is precise and worth writing down: an archive can show you that a page said something on a given date, and it can never show you that the page did not say something else in the weeks between two captures.

Can a company remove its website from the Wayback Machine?

It can ask, and the request route is a direct one to the Internet Archive rather than a technical control. This changed in 2017, when the archive said publicly that robots.txt files written for search engines work badly for archives, and stopped treating them as an instruction, having already done so for government and military sites some months earlier. Two practical consequences: pages that used to be hidden because a domain was parked have often become visible again, and a site currently blocking crawlers may still have years of history in the archive.

Is Google cache still available?

No. The cached-page links were removed in early 2024 and the cache search operator was retired shortly afterwards, with the company's search liaison explaining that the feature had been built for an era when pages routinely failed to load. Google later added a link to the Internet Archive inside its about-this-page panel, in September 2024, which is now the closest thing to the old workflow. In practice that means a single archive carries far more of this job than it did two years ago, so having a second one in your routine matters more than it used to.

Can I reuse a competitor's archived pages in my own materials?

Reading an archived copy of a page a company published to the public web is ordinary use of a public library's holdings, and the material was published deliberately in the first place. The care needed is in what you do next. An archived page is copyrighted like the original, so quote briefly and attribute rather than reproducing it wholesale. And do not present an old page as current, whether to colleagues or to a market, because a competitor's withdrawn claim from three years ago is not a statement they are making today.

Can I use an archived page as evidence?

This is what web archives are genuinely better at than anything else in competitive research. A snapshot has an address containing its own timestamp, sits on a third party's servers, and can be opened by anybody who doubts you, which is a form of provenance a screenshot in a slide deck simply does not have. Courts and regulators in several jurisdictions have accepted archived pages, usually with supporting evidence about how the archive works. For internal use the standard is lower and the principle is the same: cite the archived address, never a picture of it.

What is the difference between the Wayback Machine and archive.today?

One crawls, the other is asked. The Wayback Machine visits addresses on its own schedule and holds by far the deepest history, which makes it the right tool for the past. Archive.today captures only what somebody requested, renders the page with a full browser and stores the visual result, which makes it much better at script-heavy pages and at preserving what a page looked like rather than what was served. It also stores fewer file types. Use the first for history and the second when the first returns a broken page or nothing at all.

Why do archived pages look broken?

Because a capture stores the files a crawler could fetch, and a modern page needs far more than that. Stylesheets and scripts hosted elsewhere may not have been captured, content assembled in the browser after load was never in the response at all, fonts and images may be missing, and anything personalised or behind consent will have rendered for a crawler rather than for a visitor. The practical workaround is to read the snapshot's source as well as its rendered form, since the text you want is usually present in the file even when the page collapses.

How do I compare two versions of a competitor's page?

Use the archive's own comparison rather than reading two tabs. From a page's calendar view, open the changes option, pick two captures and it renders them side by side with added and removed text marked, which turns an hour of squinting into a minute. Start about a year apart to find the interval that contains the change, then narrow. Record both timestamps beside whatever you conclude, because the finding is only ever that something changed between two specific dates, not that it changed on one.

How do I make sure a page I found today is still provable in six months?

Save it yourself, now. The archive's save function creates a capture on request in seconds and returns a permanent public address with today's timestamp in it, and that address is the citation. Do it as a habit for competitor pricing pages, homepage claims, comparison pages that name you and anything a rival is likely to quietly revise. It costs nothing, it takes a few seconds, and it is the only part of this source you control, because everything else depends on a crawler having happened to be there.

Can I see a competitor's deleted blog posts or job adverts?

Sometimes, and the odds depend entirely on how well linked the page was. Blog posts are usually reachable through the archived index or category pages even when the individual post was never captured, since those listing pages carry the titles and dates. Job adverts are harder, because they are removed as soon as a role closes and are rarely linked from anywhere durable, so an archived careers page listing open roles is often the only surviving trace. A removed page is also a finding in itself, and worth dating.

How often does the Wayback Machine crawl a website?

There is no fixed schedule, and the frequency varies by orders of magnitude between sites and between pages on the same site. Captures come from several sources at once, including the archive's own crawls, partner crawls and pages saved on request by members of the public, which is why a page can have three captures on one day and nothing for the following two years. Rather than reasoning about the schedule, read the calendar for the address you care about, and treat its density as the ceiling on how precisely you can date any change.

Go beyond web archives on competitor pages

Flares dates every competitor change as it happens, so the history belongs to you rather than to a crawler.

Discover Flares

14-day free trial · 30-second setup