The Internet Archive is a nonprofit library that saves copies of websites, books, software, and videos so they stay accessible even after the original disappears

The Internet Archive runs the Wayback Machine, a tool that stores snapshots of web pages going back to 1996. When you visit archive.org and enter a URL, you can see what that website looked like on specific dates in the past. This matters because websites change, get deleted, or vanish entirely — a news article you read five years ago might no longer exist on the publisher's site, but the Archive often has a copy.

Beyond the Wayback Machine, the Archive also preserves millions of books (including out-of-print ones), government documents, academic papers, software programs, and television news clips. It operates physical servers in San Francisco, with backup copies stored in other locations. The organization is funded by donations, digitization services, and grants — not by selling your data or showing you ads.

Key Takeaways

  • The Wayback Machine lets you see archived versions of websites from specific dates, useful when a page has been deleted or changed.
  • The Archive preserves books, academic papers, government records, and software alongside web pages, creating a searchable library of digital history.
  • Snapshots are not automatic — the Archive crawls the web regularly but does not capture every page, and some sites block the crawler.
  • You can manually submit a URL to the Archive if you want a current snapshot saved, though it may take days or weeks to appear.
  • The Archive is free to use and does not require an account, though creating one lets you save personal collections.

How the Wayback Machine captures and stores web pages

The Wayback Machine works by sending automated crawlers across the internet to download and store copies of web pages. These crawlers visit millions of sites regularly, taking snapshots and storing them on Archive servers. When you search for a URL on archive.org, you are not seeing a live version of the site — you are viewing a saved copy from a specific date that the crawler captured.

The Archive does not capture every page on every site. It prioritizes popular websites and pages that change frequently. It also respects a file called robots.txt, which website owners can use to tell crawlers not to archive their content. Some sites, like banking portals or subscription services, actively block the Archive's crawler. If a page has never been crawled, or if the owner blocked it, no snapshot will exist.

Each snapshot includes the text, images, and basic layout from that date, though some elements may not work perfectly. Videos, interactive features, and content that required login often do not archive well. The Archive stores multiple snapshots of the same page over time, so you can see how a website evolved.

What gets preserved beyond websites

The Internet Archive's collection extends far beyond the Wayback Machine. The Books project has scanned over 20 million books, including many that are out of print and hard to find elsewhere. You can search and read these books directly on archive.org, or borrow digital copies if your library participates in the Open Library program.

The Archive also preserves government documents, academic research, court filings, and policy papers. During the 2016 U.S. election, the Archive began saving copies of government websites, concerned that some might be removed or altered. This practice continues — the Archive now maintains backups of federal agency websites and state records.

The Television News Archive stores clips from major U.S. news broadcasts going back decades. You can search by topic or date to find news coverage of specific events. The Archive also preserves software and video games, including old programs that no longer run on modern computers, through emulation projects that let you run them in your browser.

When archived pages may not exist or may be incomplete

Not every website has an archived version. If a site was brand new, very small, or actively blocked the Archive's crawler from the start, no snapshots exist. Popular sites like Google, Facebook, and Twitter have limited archives because they block or restrict crawling. News sites vary — some allow archiving, others do not.

Even when a snapshot exists, it may be incomplete. Pages that relied on JavaScript to load content often show only the basic HTML structure. Paywalled articles may appear in the archive, but the Archive respects copyright and does not preserve full text of books still under copyright protection. Images sometimes fail to load if the original server is down or the image path changed.

The most recent snapshot of a page may be weeks or months old, depending on how often the Archive crawled that site. If you need a very current version of a page that has since been deleted, you might not find it.

How to manually save a page to the Archive

You do not have to wait for the Archive's crawler to find a page. You can visit archive.org, enter a URL in the search box, and click "Save Page Now." This tells the Archive to create a snapshot immediately. The page will be queued for capture, though it may take anywhere from a few minutes to several days to appear, depending on the Archive's workload.

Manual saves are useful when you want to preserve something you found online before it disappears. Journalists, researchers, and activists often use this feature to create timestamped records of web pages. Once saved, the snapshot is public and searchable on archive.org.

If you create a free account on archive.org, you can build personal collections of saved pages and organize them by topic. This does not change how the Archive works — it just gives you a private dashboard to track pages you have saved.

Privacy and copyright considerations when using the Archive

The Internet Archive respects copyright law. It does not preserve the full text of books still under copyright protection, though you can often read excerpts or borrow digital copies through partner libraries. Out-of-print books and older works are more likely to be fully available.

Archived web pages are public and searchable. If you posted something online years ago and the Archive captured it, that snapshot is now part of a searchable historical record. You can request removal of specific pages by contacting the Archive, though the process is not automatic. The Archive will remove pages if the copyright holder requests it, or if the content violates certain policies.

The Archive itself does not track who views archived pages. It does not use cookies to follow you across the web, and it does not sell data. However, like any website, archive.org logs basic information about visits (IP address, browser type) for server maintenance.

Alternatives and related tools for saving web content

If the Wayback Machine does not have what you need, other options exist. Perplexity and some other search engines create snapshots of pages they index. Browser extensions like Evernote Web Clipper and Notion Web Clipper let you save pages to your personal account. These tools store copies on their own servers rather than in a public archive.

For academic research, Google Scholar and institutional repositories preserve published papers. For news articles, some publishers maintain their own archives, and services like NewsGuard track which outlets have reliable archives. For social media posts, specialized tools like the Wayback Machine for Twitter (now X) exist, though they have limitations.

If you want to preserve something important, using multiple methods is wise. Save a copy to the Internet Archive, download a PDF to your computer, and consider storing it in cloud storage. This redundancy protects against any single service going down or removing content.

Frequently Asked Questions

Can I remove my website from the Internet Archive?

Yes. You can add a line to your site's robots.txt file to block future crawling, and you can request removal of existing snapshots through archive.org's exclusion request form. The Archive will honor these requests, though it may take time to process. Note that this only affects the Internet Archive — other services may have their own copies.

Is it legal to use archived versions of copyrighted content?

The Archive operates under fair use and copyright law. Archived news articles and web pages are generally treated the same as the originals — you can read them for research or reference, but republishing them without permission is not legal. Books under copyright are restricted to preview or lending through partner libraries.

Why does an archived page look broken or incomplete?

Archived pages often lack interactive features, videos, or content that required JavaScript to load. Images may not display if the original server is down. The Archive captures what it can, but complex modern websites do not always archive perfectly. Older snapshots are more likely to be incomplete than recent ones.

Can I download an entire archived website?

The Archive does not provide a bulk download tool for most sites, but you can download individual pages or use third-party tools like Wget to grab multiple pages. For books, you can download PDFs or ePub files directly from archive.org. Check the Archive's terms of service and copyright status before downloading large amounts of content.

How far back does the Wayback Machine go?

The Wayback Machine has snapshots dating back to 1996, though coverage is sparse for the earliest years. Most major websites have regular snapshots from the early 2000s onward. The further back you go, the fewer snapshots exist and the more likely pages are incomplete.