From shuttered airline inboxes to pallets of out-of-print paperbacks, artificial intelligence companies are racing to buy up forgotten data troves, turning corporate detritus and old books into fuel for the next generation of powerful models.

Get the latest news straight to your inbox!

AI’s Data Hunger Reaches Dead Airlines and Dusty Bookshelves

Google’s Bid for a Dead Airline’s Digital Remains

A recent bankruptcy auction for Spirit Airlines’ technology assets offered a glimpse into how aggressively major tech firms are pursuing proprietary datasets. Court filings indicate that Google agreed to pay about 10 million dollars for a vast slice of the defunct carrier’s digital history, outbidding at least one specialist AI firm for the rights.

According to published coverage, the package includes roughly 100 million employee emails, hundreds of millions of Microsoft Teams chats, extensive software repositories and internal business systems data. The information is described in legal documents as “de-identified,” with passenger profiles and loyalty records excluded, but privacy advocates note that corporate communications often contain personal and commercially sensitive details even after names and obvious identifiers are removed.

Publicly available information on the deal suggests Google intends to use the data to improve its AI models and productivity tools. The purchase illustrates how archives that once would have been treated as routine back-office records are now being reconceived as strategic assets, potentially worth millions once repurposed as training material for large language models and other systems.

For the travel industry, the auction underscores an emerging reality: when a carrier disappears, its digital footprint may live on inside AI products, raising new questions for workers, customers and regulators about who controls historical data and how it is reused.

Corporate Exhaust Becomes Premium AI Training Fuel

The Spirit transaction highlights a broader shift in how bankrupt or struggling companies are valued. Alongside aircraft, routes and brand names, distressed firms now hold a less visible asset that technology buyers prize: years of internal conversations, operational logs and code repositories that capture how real businesses actually function.

Specialist intermediaries have begun to appear, offering to help failed start-ups package their source code, chat histories and email archives for sale to AI developers. Reports indicate that these marketplaces pitch so-called “organizational exhaust” as uniquely rich material for training models that can write software, summarize meetings or mimic white-collar workflows.

Travel companies are part of this trend because they generate dense streams of structured and unstructured information, from scheduling and crew coordination to dynamic fare-setting and customer service correspondence. Even when customer identities are removed, the resulting data can reveal complex decision patterns that AI systems may learn to replicate.

Privacy researchers point out that de-identification is not a guarantee that individuals cannot be re-identified, particularly when large datasets are combined. At the same time, there is little harmonized regulation governing how long companies may retain such archives or what happens to them when a business shuts down, leaving workers and travelers with limited visibility into where their historical interactions could resurface.

Old Books Get a Second Life as Training Data

While one wave of AI data collection sweeps through abandoned corporate servers, another is hitting the stacks of used book warehouses and distributors. Legal filings and investigative reporting describe how major AI developers have spent tens of millions of dollars acquiring physical books in bulk, then cutting off the bindings so the pages can be fed through high-speed scanners.

In some cases, according to court documents and media coverage, the volumes are destroyed after scanning, turning entire print runs into one-time inputs for training models. Travel and geography titles, guidebooks and narrative nonfiction are attractive targets because they contain detailed descriptions of destinations, cultures and historical events that can enrich a system’s responses to travel-related prompts.

This industrial-scale digitization goes beyond earlier library and search engine projects that focused on discoverability. Instead, the goal is to convert as much high-quality text as possible into machine-readable form, then internalize it inside proprietary AI models. That shift has heightened concern among authors, publishers and librarians who see unique or out-of-print works being pulped for datasets that remain closed to the public.

Observers note that physical books, unlike pirated files, can be lawfully purchased at scale, potentially offering AI firms a stronger legal footing. Recent rulings in U.S. courts have treated some large-scale book scanning for AI training as permissible fair use, encouraging companies to expand these programs even as parallel copyright cases continue.

The rush to feed AI systems with text from books, news archives and online repositories has triggered a cascade of lawsuits around the world. Authors, record labels and news organizations have filed complaints alleging that their works were copied without consent into massive datasets, including collections built from pirated ebooks and so-called “shadow libraries.”

High-profile cases have named leading technology firms and cited datasets such as Books3, a compilation of nearly 200,000 books that plaintiffs say was assembled from unauthorized copies. Courts are now grappling with whether transforming copyrighted works into training data is a fair use, and how to weigh the public benefits of improved AI tools against the potential loss of control and income for creators.

Recent decisions have begun to sketch a mixed landscape. Some judges have accepted arguments that training on large text corpora is transformative and therefore lawful, while others have expressed concern about the scale and secrecy of data acquisition. Legal analysts suggest that the emerging case law may treat scanning lawfully acquired physical books differently from harvesting files from pirate sites, incentivizing the kind of bulk print purchases now alarming booksellers.

For the travel sector, the outcome will shape how everything from guidebooks and destination histories to airline manuals and route maps may be repurposed in future AI tools. If courts continue to green-light broad scraping of such material, much of the industry’s accumulated knowledge could be ingested into systems that then compete with traditional publishers and information providers.

Privacy Questions Trail AI Into the Inbox and the Cabin

Public debate has largely focused on copyright, but the Spirit Airlines data sale highlights a parallel set of privacy issues. Corporate email and chat archives often contain personal information about employees and travelers, even if direct identifiers are later removed. That makes them especially sensitive when repurposed for training systems that can reproduce patterns of language or internal decision-making.

Major AI developers have published assurances that personal emails and private documents are not used indiscriminately to train their largest models, or that such use is subject to explicit consent and opt-out settings. At the same time, policy documents acknowledge that some user-generated content, product telemetry and connected app data may be drawn into anonymized training pipelines, with varying controls across jurisdictions.

The lack of comprehensive federal privacy legislation in the United States means that protections can depend heavily on individual company policies and terms of service. Researchers warn that, as more distressed firms seek to monetize historical data, employees and customers may have little practical way to prevent their old messages, booking patterns or complaint letters from helping teach AI systems how airlines and other travel businesses operate.

Advocates for stronger safeguards argue that data involving travel itineraries, location histories and identity documents deserves heightened treatment, regardless of whether it sits in an active reservation system or in a decade-old email archive. They are calling for clearer rules on when such information can be sold, how it must be anonymized and whether people should have a right to demand its deletion before it can be transformed into AI training fuel.