Artificial intelligence developers are turning to unconventional troves of information, from the internal emails of defunct airlines to pallets of old and rare books, as they compete for the vast quantities of data needed to train next-generation models, raising urgent questions about privacy, ownership and the future of cultural heritage.

Get the latest news straight to your inbox!

AI Firms Tap Defunct Airlines and Old Books for Training Data

Dead Airlines’ Digital Remains Become AI Training Gold

In the latest sign of how valuable proprietary data has become, recent bankruptcy filings show technology companies vying for access to the internal systems of failed travel firms. Court documents from the wind-down of a major U.S. low-cost carrier indicate that Google agreed to pay around 10 million dollars for a package of corporate information that includes roughly 100 million employee emails, hundreds of millions of collaboration messages and extensive operational records, with the data earmarked for use in improving artificial intelligence models and products.

The auction attracted bids from specialist AI firms as well as large platforms, reflecting how even the digital exhaust of a collapsed airline has become a strategic asset. While the data is described as de-identified and stripped of passenger records, it encompasses years of internal communications, software repositories and pricing databases that could help refine language models, code-generation tools and forecasting systems.

Legal experts note that selling corporate archives in bankruptcy is not new, but the explicit framing of such material as training fuel for AI systems marks a shift. The case underscores how distressed travel and logistics companies, rich in operational data but short on cash, are emerging as unexpected suppliers in the AI supply chain.

Privacy advocates argue that the use of workplace communications in this way highlights a growing gap between employee expectations and the afterlife of their data. Staff whose emails and chat logs now sit inside AI training corpora may never have imagined their day-to-day messages would outlive the airline itself and be repurposed to power future software.

Old Books, New Models: How Physical Libraries Are Being Mined

At the same time, reports from booksellers in Europe, North America and Australia describe unusually large bulk purchases of used, rare and out-of-print titles by intermediaries linked to the AI sector. Investigations by technology and book-industry media indicate that many of these volumes are shipped to specialized facilities where bindings are removed, pages are scanned using high-speed equipment, and the physical books are then pulped or recycled once the text has been captured.

Secondhand bookshops recount orders for thousands of niche, technical or regional works that had sat unsold for years, followed by the realization that they were unlikely to see the physical copies again. For AI developers, such books are attractive precisely because they are not widely available online. Scanning them offers access to specialized knowledge that can improve the breadth and factual coverage of large language models.

The practice has opened a new front in the copyright debate surrounding generative AI. Major academic and trade publishers in the United States and Europe have launched lawsuits accusing Google and other firms of using millions of books and journal articles to train systems like Gemini without proper authorization, arguing that texts supplied for search or sales purposes were never licensed for model training. Industry groups characterize the mass digitization of physical and out-of-print works for AI as a form of unlicensed extraction that goes well beyond earlier library-scanning projects.

AI companies, for their part, often cite legal doctrines such as fair use in the United States or text-and-data-mining exceptions in parts of Europe, and point to past cases involving search engines and digital libraries. But critics say the scale, commercial orientation and irreversible destruction of physical copies in some AI scanning pipelines make this a qualitatively different enterprise from previous efforts to make books searchable.

The scramble for data is unfolding against a backdrop of intensifying litigation and regulatory review. In U.S. courts, authors’ groups, music and image rights holders, and now large book and educational publishers have filed a series of class actions against leading AI companies. The complaints generally allege that developers copied vast numbers of copyrighted works into training datasets without consent or compensation, and that some models can reproduce protected material too closely for the practice to qualify as lawful.

Recent rulings have begun to sketch the outlines of how judges may treat AI training under copyright law, with some decisions finding that certain forms of training can be considered transformative, and others rejecting broad immunity claims and stressing that there is no blanket exemption for machine learning. Settlements have already produced multibillion-dollar payouts in at least one high-profile author case, signaling both the financial stakes and the willingness of some firms to resolve disputes rather than risk precedent-setting trials.

Regulators are also taking notice. Policy papers from agencies in the United States and Europe describe generative AI training as heavily reliant on books, academic papers, legal documents and other professionally produced works, and warn that unlicensed extraction could undermine the economic foundations of publishing and journalism. Proposals range from clearer opt-out mechanisms for rightsholders to new collective licensing schemes that would allow AI developers to pay for access to large bodies of text under standardized terms.

Travel and hospitality companies are watching these developments closely, as many rely on partnerships with technology providers that incorporate AI tools. The way courts and regulators define permissible data use in training may shape how airlines, hotels and online booking platforms manage their own archives, loyalty programs and customer communications.

What It Means for Travelers and Digital Privacy

For individual travelers, the prospect of AI models trained on the remnants of airlines and on digitized books may feel remote, yet the implications are tangible. Corporate datasets built from years of flight schedules, disruption logs and customer support interactions can help AI systems forecast delays, optimize pricing and improve automated service agents that passengers increasingly encounter when rebooking or filing complaints.

However, the same datasets raise questions about long-term stewardship of personal and behavioral information. Even when consumer records are removed from bankruptcy sales, operational systems often contain indirect traces of traveler behavior that can be valuable for training recommendation engines and risk models. Privacy advocates warn that as distressed companies seek to monetize every possible asset, the line between acceptable and intrusive reuse of such information may blur.

The book-scanning controversies also speak to a broader concern about cultural preservation. When rare or regionally significant titles are destroyed after digitization, critics argue that physical artifacts that once anchored local memory and scholarship are being converted into proprietary model weights controlled by a handful of firms. For travelers who seek out independent bookshops, archives and cultural institutions, the idea that these spaces have become quiet feeders for opaque AI pipelines has proved unsettling.

Consumer groups are urging clearer disclosure from both AI vendors and the companies that supply them with data, including airlines and travel intermediaries. They argue that passengers and readers should at least be informed when their interactions or purchased materials are likely to end up inside training datasets, even in anonymized form. Transparency, they contend, is a prerequisite for meaningful consent and for any future revenue-sharing arrangements.

A Global Data Supply Chain With Local Flashpoints

Taken together, the auction of a dead airline’s digital archives and the steady flow of used books into scanning centers illustrate how generative AI has created a new kind of global data supply chain. Bankruptcy courts, secondhand warehouses and logistics hubs have become upstream nodes feeding the models that increasingly shape search results, travel planning tools and automated customer service across the tourism sector.

Local flashpoints are emerging wherever this supply chain intersects with communities. In some cities, booksellers report feeling complicit in the disappearance of their own cultural heritage when they fulfill bulk orders destined for shredders. In others, former airline employees express discomfort at learning that their internal chats and code repositories may now be part of commercial AI systems, even if scrubbed of direct identifiers.

Industry analysts say these tensions are unlikely to slow the hunt for data in the short term. As generative models grow more sophisticated, developers need not only larger volumes of text, but also more varied, domain-specific material, from aircraft maintenance manuals to regional histories. That demand gives owners of obscure archives, including in the travel and cultural sectors, new bargaining power, but also exposes them to complex ethical and legal choices.

For travelers and readers, the outcome of the current legal cases and policy debates will help determine whether the stories, records and everyday communications of the past remain accessible public resources or disappear into closed commercial systems. The way AI companies handle the digital remains of dead airlines and the pages of aging books may become an early test of how the technology industry balances innovation with respect for privacy and cultural memory.