Jim Tinsley compiles the definitive archival FAQ for Project Gutenberg, capturing the collaborative mechanics behind the world's oldest digital library as it operated in two thousand and two. The text addresses the three primary production bottlenecks faced by early volunteers: acquiring eligible public domain books, securing scanners or typing skills, and investing the forty painstaking hours typically required to produce a single e-text. Through detailed sections covering copyright clearance, scanning settings, proofreading strategies, and file formatting standards, the work demystifies the transition of paper volumes into plain vanilla ASCII and HTML archives.
The author details how Michael Hart founded the project in nineteen seventy-one with a digital copy of the Declaration of Independence, eventually scaling production targets to two hundred books per month by mid-two thousand and two. The text outlines the exact five steps of e-text creation, emphasizing that Project Gutenberg relies entirely on decentralized self-assigned tasks rather than a central command structure. Volunteers are guided through finding pre-nineteen twenty-three texts via online booksellers, submitting title page and verso images for copyright clearance, and utilizing distributed proofreading frameworks.
Special attention is given to the technical rationale behind plain text formatting, explaining that ASCII ensures texts remain readable across centuries despite rapidly changing technological landscapes. The manual addresses common obstacles such as OCR scanning errors, handling British pound signs and Greek characters, resolving line-wrapping issues, and submitting files via FTP or automated web uploads. It also highlights the crucial role of community helpers, mirrors on five continents, and dedicated mailing lists in sustaining the archive without relying on commercial digital rights management or tracking user downloads.