A plain-text book can look like the least sophisticated version of an ebook: no typeface, no page design, no illustrations, often not even italics. This apparent crudeness is one reason the format has lasted so well.
Project Gutenberg has been making ebooks since 1971, through enough generations of computing to watch many once-current formats become obsolete. Since 2004, almost every ebook it releases has plain-text and HTML master versions. Its explanation is practical. Open, editable formats can be read without a vendor's software, corrected without rebuilding a proprietary package, and converted into whatever a future reader needs. A .txt file from decades ago remains remarkably easy to open. Project Gutenberg, “File Formats Utilized by Project Gutenberg”
Part of that durability comes from keeping the representation simple.
The useful question is not whether a plain-text file is a good or bad copy of a book. It is which parts of the book the file was built to carry.
The words survive more easily than the edition
Library catalogs have long needed ways to distinguish an intellectual work from its particular forms. The Library of Congress's BIBFRAME model separates a Work, representing the underlying intellectual creation, from an Instance, a particular published embodiment, and an Item, an actual copy of that instance. Library of Congress, “Overview of the BIBFRAME 2.0 Model”
The distinction is useful here even though a Project Gutenberg file does not map neatly onto one BIBFRAME category.
Take Pride and Prejudice. The sequence of Austen's sentences can travel astonishingly well into plain text. A particular printed edition carries more with it: the publisher and date, pagination, spelling and editorial choices, type design, chapter openings, perhaps an introduction or notes, perhaps illustrations. One individual copy may carry still more: an inscription, a bookplate, pencil marks, a repaired spine, evidence of who owned it and how it was used.
When that book becomes .txt, much of the wording can survive while evidence belonging to that edition and copy falls away.
For ordinary reading, this can be exactly the right bargain. If the question is “What happens in chapter 34?” or “Where does this phrase occur?”, preserving the words and making them searchable may matter far more than preserving the width of the margins. If the question is “What did the 1894 edition ask its reader to see on this page?”, plain text is suddenly a poor witness.
Typography is sometimes information
A text file can contain letters, punctuation, spaces, and line breaks. It does not natively know that a word appeared in italics, that a heading used small capitals, that a printer letterspaced a name, or that one passage was centered while another was indented.
Plain-text editions often invent conventions to compensate. _A word_ may stand for italics. Capital letters may substitute for small caps. Blank lines can suggest larger visual divisions. These conventions are useful, but they are descriptions or replacements for typography rather than the typography itself.
How much that matters depends on the book.
In a conventional prose novel, the exact body type may contribute to the edition's character without changing the basic intelligibility of the sentences. In a concrete poem, a typographically experimental novel, a dictionary, a mathematical work, or a book whose hierarchy depends heavily on size and placement, the visual system can carry information that refuses to collapse cleanly into a character stream.
The Library of Congress treats this as a real preservation question. Its guidance for textual content separates the integrity of a document's logical structure from the integrity of its layout, fonts, design, diagrams, mathematics, and pagination. It explicitly notes that no single digital format is best for every textual object. Library of Congress, “Quality and Functionality Factors for Textual Content”
This is a better framework than saying “formatting gets lost.” Sometimes formatting is decoration. Sometimes it tells the reader what kind of thing a passage is, where it belongs, or how it relates to something beside it.
A page contains relationships that a line of text does not
Printed pages are spatial arrangements. A note may sit beneath the sentence it qualifies. An illustration may face the paragraph it depicts. Two columns may be meant to be compared. A running head can tell you where you are. A table depends on alignment. Marginal notes live beside rather than after the main text.
A .txt file has to turn that surface into an order. Something must come first, then second, then third.
That can be harmless. A chapter heading followed by three paragraphs has an obvious linear sequence. Other pages force editorial decisions. Does a footnote appear immediately after its marker, at the end of the paragraph, or at the end of the chapter? Does a caption come before or after the place where the illustration used to be? How do you represent two columns whose relationship is horizontal rather than sequential?
Pagination creates a related problem. Page numbers can be inserted into plain text as markers, but the file no longer contains the page those numbers refer to. For a casual reader this may be irrelevant. For citation, textual scholarship, or comparison of editions, it can be decisive.
Project Gutenberg's collection-development policy is unusually candid about this. Its ebooks are digital editions derived from printed works, not promises of exact facsimile reproduction. Production can include removing running headers and footers, joining words broken across printed lines, relocating notes, and making choices about page numbers and illustrations. The resulting file is often more convenient to read and search precisely because it has stopped imitating the printed page. Project Gutenberg, “Collection Development Policy”
Pictures do not survive merely because the words around them do
Illustrations are the obvious loss, but not the only one. Maps, diagrams, musical notation, printer's ornaments, tables, mathematical display, decorative initials, and other non-textual elements all require some representation beyond ordinary plain text.
Project Gutenberg commonly pairs plain text with HTML, which can retain images and richer structure. It also uses formats such as LaTeX when mathematical notation requires more than a simple character stream. Its preservation practice is therefore not “plain text is enough for every book.” It is closer to “plain text is worth having whenever the work can sensibly support it.” Project Gutenberg, “File Formats Utilized by Project Gutenberg”
That distinction matters for books whose pictures are not supplementary. An illustrated natural history, an architectural treatise, a children's picture book, or an edition built around an artist's images cannot be reduced to prose without changing what the object is doing.
The individual copy disappears almost completely
Digitization discussions often stop at typography and images, but a book historian may care about things that were never part of the publisher's ideal text at all.
A reader writes a disagreement in the margin. A family records births on a blank leaf. A library stamps its name over an earlier owner's bookplate. Someone has a cheap edition rebound in leather. A bookseller's ticket remains inside the cover. Pages are cut, repaired, stained, censored, or left unopened.
These details belong to an individual object. They can reveal ownership, reading, circulation, taste, institutional history, and sometimes the circumstances under which a text mattered.
A clean transcription quite properly throws most of this away. That is not a defect if its purpose is to let thousands of people read the text. It becomes a defect only when we forget that the cleaned text and the particular book answer different questions.
Digital does not require flattening everything
The alternatives are more interesting than “physical book versus text file.” Digital editions can preserve different layers with different techniques.
HTML and EPUB can mark headings, quotations, notes, lists, links, emphasis, images, and other structural features while allowing the display to reflow for different screens. A PDF or page-image facsimile can preserve a particular visual arrangement. Scholarly encoding systems such as the Text Encoding Initiative can go considerably further, recording logical structure, aspects of rendition, revisions, and relationships between transcription and page images.
The current TEI Guidelines even distinguish between encoding a text according to its logical structure and representing a source as a two-dimensional written surface. A digital facsimile can be linked to transcription so that the searchable words and the visible page remain connected. Text Encoding Initiative, facsimile examples from the TEI P5 Guidelines
This is labor-intensive, which is why no sensible project encodes every surviving feature of every book. Preservation is always selective. The important part is knowing what the selection is.
And digital text is not immaterial. A .txt file is itself a new artifact with an encoding, lineation, editorial history, filesystem history, and software environment. Digitization does not remove materiality so much as create a different kind of material record with different evidence.
Plain text gains something by losing so much
A facsimile can show you a page that plain text cannot. It is also harder to search as text unless OCR or transcription has been added. A richly structured digital edition can retain distinctions that .txt cannot, but it also depends on more markup, more specifications, and more software that knows what to do with them.
Plain text is radically portable. It can be searched, copied, diffed, indexed, analyzed in a corpus, read by simple software, transformed into new formats, and stored at trivial cost. It is unusually independent of the interface through which you happen to encounter it. Project Gutenberg's bet on plain text makes more sense after decades of software churn than it did at the beginning.
The Library of Congress's Recommended Formats Statement captures the tradeoff from another direction. For digital books it prefers structured markup such as EPUB3 or documented XML formats, and high-quality page-layout formats where appropriate; plain text remains acceptable farther down the hierarchy. Preservation value is not one-dimensional. Openness and longevity matter, but so do structure, rendering, and significant features of the work. Library of Congress, “Recommended Formats Statement: Textual Works”
So the right format depends on what you want to keep.
If you need the language of a conventional prose work in a durable, searchable form, .txt can be magnificent. If you need the logical structure of a complex publication, use a format that can name that structure. If you need the appearance of a particular edition, keep page images or a format designed to preserve layout. If you need to study a specific copy, no transcription can substitute for documentation of the object, and sometimes for the object itself.
A book contains several kinds of evidence at once. Plain text survives by choosing one of them and carrying it very lightly.
Sources
- Project Gutenberg, “File Formats Utilized by Project Gutenberg”
- Project Gutenberg, “Collection Development Policy”
- Library of Congress, “Quality and Functionality Factors for Textual Content”
- Library of Congress, “Recommended Formats Statement: Textual Works”
- Library of Congress, “Overview of the BIBFRAME 2.0 Model”
- Text Encoding Initiative, TEI P5 Guidelines for Electronic Text Encoding and Interchange, facsimile examples
A relevant Ulix tool
Guten
Guten is a free Android reader for 70,000+ public-domain books from Project Gutenberg. It takes portable digital texts and gives them a reader designed for actual use: offline reading, search, themes, highlights, notes, and bookmarks. The underlying lesson of plain text still matters here: durable access begins with texts that are not locked to one storefront, account, or device.