The Internet Archive is a nonprofit library that stores copies of websites, books, and media so you can see what pages looked like in the past
The Internet Archive, also called the Wayback Machine, takes snapshots of websites on specific dates and saves them. When you visit a URL on the Wayback Machine, you can see what that page looked like years ago — the text, the layout, the images, everything as it appeared on that date. The Archive also stores millions of books, audio files, and videos in one searchable place.
The organization has been running since 1996 and is funded by donations and grants, not by advertising or selling your data. You can use it for free. The main reasons people use it: to find a page that no longer exists, to see how a website has changed over time, to read books that are out of print, or to verify what something said on a specific date.
Key Takeaways
- The Wayback Machine at archive.org lets you type in any website URL and see saved versions from different dates going back to the mid-1990s.
- Not every page is captured on every date — the Archive crawls the web automatically, so coverage varies by site and time period.
- You can read millions of books, listen to audio recordings, and watch videos directly on the Archive's site without paying.
- The Archive does not store passwords, login information, or pages behind paywalls, so some content will never be captured.
How the Wayback Machine works and what you can find there
Go to archive.org and type a website URL into the search box. The Wayback Machine shows you a calendar with blue dots on dates when that page was captured. Click any date and you see the page as it looked on that day. You can jump between years or months to watch how a site changed — a company's homepage redesign, a news article before it was edited, a product page before the price went up.
The captures are not complete. The Archive's web crawlers automatically visit millions of sites, but they do not capture every page or every update. A small personal blog might have only a handful of snapshots. A major news site might have dozens per year. Pages that require you to log in are not captured, and neither are pages behind paywalls or password-protected areas. If a website owner asks the Archive to remove their site, those snapshots disappear.
The oldest captures go back to 1996, but most sites have no snapshots before 2000 or 2005. The further back you go, the fewer snapshots exist. If you are looking for a page from a specific date and the calendar shows no dot on that date, try dates nearby — you might find a version from a few days or weeks earlier.
Reading books and accessing media on the Archive
The Internet Archive's Open Library section holds millions of books you can read online for free. You can search by title, author, or ISBN. Some books are in the public domain — meaning copyright has expired — and you can read them in full. Others are still under copyright, but the Archive lets you borrow them for 14 days at a time, the way a library does. You do not need a library card.
Beyond books, the Archive stores audio recordings, including music, podcasts, and spoken-word performances. It also holds millions of videos, from old television broadcasts to home movies to educational content. You can stream most of this directly from the site. The Archive also preserves software and video games from decades past, which you can run in your browser.
Why the Internet Archive matters for digital safety and record-keeping
The Internet Archive serves as a backup for the internet itself. Websites disappear — companies shut down, pages get deleted, URLs break. Once something is gone from the live web, it is often gone forever unless the Archive captured it. This matters if you need to prove what a website said on a certain date, or if you are researching a company's history and want to see their old claims.
For your own digital safety, the Archive can help you verify information. If someone claims a news article said something, you can check the Wayback Machine to see what the article actually said on the date in question. If a website changed its terms of service or privacy policy, you can see the old version. This is one reason to keep your own backups of important pages — the Archive is a safety net, not a may provide that everything you care about will be there.
What the Internet Archive does not capture and why
The Archive does not store anything behind a login. If a page requires a username and password, the crawlers cannot see it, so it is never captured. This includes your email, your social media accounts, your banking site, and any private documents. The Archive also respects when website owners use a robots.txt file to tell crawlers to stay away — many sites opt out of being archived.
Pages that load content dynamically — meaning the page builds itself with JavaScript after you visit — are harder to capture accurately. The Archive's crawlers may see a blank page or an incomplete version. Very new websites might not have been crawled yet, so recent pages may not appear in the Wayback Machine for weeks or months.
The Archive also does not store copyrighted material that is still actively sold or distributed, with limited exceptions for preservation purposes. If a book is in print and the publisher is selling it, you will not find the full text on Open Library — you will see metadata and a link to borrow it instead.
How to use the Internet Archive to save your own content
You can manually request that the Archive capture a page right now. Go to archive.org, type the URL into the search box, and look for a button that says "Save Page Now" or similar. Click it and the Archive will crawl that page and add it to its collection. This is useful if you want to preserve something you think might disappear — a news article, a product listing, a blog post.
The Archive also has a browser extension for Chrome and Firefox that adds a button to your toolbar. Click it on any page and you can save that page to the Archive when ready. This is faster than going to the website each time. Keep in mind that the Archive's crawlers may not capture everything perfectly — images might not load, styling might be off, or interactive elements might not work. But the text and basic structure will be there.
If you want to save pages for your own personal use across devices, the Internet Archive is one option, but it is not private — anything you save there is public and searchable. For truly private backups, use a tool designed for that purpose, like a note-taking app, a cloud storage service, or a browser bookmark manager that syncs across your devices.
Privacy and security considerations when using the Archive
The Internet Archive is a public library. Anything you search for there, anyone else can search for too. If you look up a website on the Wayback Machine, that search is not private. The Archive itself does not track you or sell your data, but your internet service provider can see that you visited archive.org, just as they can see any website you visit.
Be careful what you assume about accuracy. Captured pages are snapshots, not perfect replicas. Links might be broken, images might not load, and the layout might look wrong. Do not rely on the Archive as your only source of truth for something important — use it as one piece of evidence alongside other sources.
If you are saving sensitive information to the Archive, remember that it is permanent and public. Do not save pages that contain your personal details, financial information, or anything you would not want anyone to find. The Archive does not delete pages easily, even if you ask.
Frequently Asked Questions
Can I remove my website from the Internet Archive?
Yes. Website owners can use a robots.txt file to tell the Archive not to crawl their site going forward, or they can contact the Archive directly to request removal of existing snapshots. The process takes time, and old snapshots may remain visible for a while. If you own a site and want it removed, visit archive.org/about/exclude.php for instructions.
Is it legal to use content I find on the Internet Archive?
It depends on the copyright status. Public domain works are free to use. Books you borrow from Open Library are for personal reading only, not for republishing. Archived web pages are generally for research and reference. When in doubt, check the Archive's terms of service or contact the copyright holder before reusing anything.
Why is a page I am looking for not on the Wayback Machine?
The site might have blocked the Archive's crawlers, the page might be too new, or it might have been removed at the owner's request. Try searching for a different date or a related page on the same domain. If the page was behind a login or paywall, it was never captured.
Can I read books from Open Library?
Yes, for many books. Public domain books can usually be downloaded as PDF or EPUB files. Borrowed books can sometimes be downloaded depending on the format and the publisher's permissions. Look for a read button on the book's page.
Does the Internet Archive have everything that was ever on the internet?
No. The Archive captures a large portion of the public web, but not all of it. Private sites, password-protected pages, and sites that opted out are not there. Coverage is spotty for small sites and very old pages. It is a valuable resource, but not a complete record.