LCSH.info All articles
Technology & Innovation

The Future of Finding: How US Libraries Are Reimagining Subject Organization in the Age of Linked Data

LCSH.info
The Future of Finding: How US Libraries Are Reimagining Subject Organization in the Age of Linked Data

Photo: Warner, Charles Dudley, 1829-1900; Mabie, Hamilton Wright, 1846-1916; Runkle, Lucia Isabella (Gilbert), 1844-, No restrictions, via Wikimedia Commons

For most of the twentieth century, if you wanted to find a book on a specific topic in an American library, you consulted the subject catalog and hoped that the term you had in mind matched the one a cataloger had chosen, possibly decades earlier, to describe that concept. This gap between the language of searchers and the language of catalogers has always been one of the fundamental tensions in library science. Today, a convergence of new technologies and shifting professional priorities is pushing US libraries to rethink how subject organization works—and whether LCSH, for all its authority and scale, is sufficient to meet the demands of twenty-first-century information access.

The Limits of a Controlled Vocabulary

LCSH remains the most widely used subject vocabulary in American academic and public libraries. Its depth is extraordinary: more than 340,000 authorized headings covering virtually every domain of human knowledge, maintained by the Library of Congress and updated continuously. For traditional catalog searching, it provides a degree of precision and consistency that free-text keyword searching cannot match.

But controlled vocabularies have inherent constraints. They require catalogers to translate a work's subject matter into an approved term, and that translation is only useful if the searcher knows to use the same term. When a patron searches for "climate change" and the catalog record uses "climatic changes"—the authorized LCSH form until relatively recently—relevant materials may never surface. Multiply this friction across thousands of topics and millions of users, and the cumulative effect on access is substantial.

Beyond terminological mismatch, LCSH faces criticism for the time and expertise required to apply it correctly. Professional-level cataloging is resource-intensive, and many smaller libraries, special collections, and digital repositories lack the capacity to provide thorough subject analysis for every item in their holdings. The result is uneven coverage that disadvantages precisely the collections—community archives, Indigenous knowledge repositories, local history collections—that often document the most underrepresented communities.

Enter Linked Data: A Different Model for Description

The linked data movement, which gained significant momentum in library circles following the Library of Congress's own Bibliographic Framework Initiative (BIBFRAME), offers a fundamentally different approach to knowledge organization. Rather than assigning a string of text from a controlled list, linked data describes resources using structured relationships between machine-readable identifiers. A book about a specific person, place, or concept is linked directly to authoritative data about that entity—its relationships to other entities, its alternate names, its context within a broader knowledge graph.

For subject access, this has transformative implications. A linked data environment can expose connections that a traditional catalog record cannot. A researcher interested in the economic history of Chicago's meatpacking industry might, in a linked data system, move fluidly between the corporate histories of specific companies, the labor organizing histories of the workers who staffed them, the public health literature generated by the conditions within them, and the literary works—Upton Sinclair's The Jungle among them—that documented their social impact. These connections exist in the traditional catalog only if catalogers made them explicit through carefully chosen headings; in a linked data environment, they can emerge from the structure of the data itself.

The Library of Congress has been developing BIBFRAME as a replacement for the MARC record format that has underpinned library catalogs since the 1960s. BIBFRAME is designed to make library data interoperable with the broader web of linked open data, including resources such as Wikidata, the Getty Vocabularies, and DBpedia. Several major institutions—among them Harvard, Cornell, and the University of Michigan—have been piloting BIBFRAME-based workflows, and the findings from these projects are shaping how the broader profession thinks about the transition.

AI-Assisted Cataloging: Promise and Caution

Alongside linked data, artificial intelligence tools are entering the cataloging workflow with increasing frequency. Machine learning models trained on large collections of catalog records can now suggest subject headings, generate metadata for digital images, and flag potential inconsistencies in existing records—tasks that previously required hours of human attention.

Vendors including OCLC, Ex Libris, and a growing number of specialized startups are offering AI-assisted cataloging modules integrated into existing library management systems. Early adopters report significant time savings, particularly for high-volume digital projects where professional cataloging of every item would be impractical. The Biodiversity Heritage Library, for instance, has used machine learning to generate subject metadata for thousands of digitized natural history texts that might otherwise have remained effectively undiscoverable.

However, AI tools introduce their own risks. Models trained on historical catalog data will reproduce the biases embedded in that data—including the very LCSH problems discussed above. If a training set consistently associates certain communities with certain headings, an AI system will learn and perpetuate those associations. Librarians working with these tools emphasize that human review remains essential, particularly for materials related to marginalized communities, contested historical events, or emerging fields where the vocabulary itself is still being negotiated.

Community Vocabularies and the Case for Pluralism

A third strand of innovation involves the development of community-specific vocabularies that either supplement or, in some cases, replace LCSH for particular collections. The Homosaurus, an international linked data vocabulary for LGBTQ+ concepts, provides more granular and affirming terminology than LCSH offers for queer subjects. The Mukurtu platform, developed in collaboration with Indigenous communities, allows groups to apply their own metadata frameworks to digital cultural heritage materials, controlling both how items are described and who can access them.

These community-driven approaches reflect a broader shift in how the profession thinks about the purpose of subject description. Rather than a single authoritative vocabulary applied uniformly across all collections, an emerging consensus holds that different communities may need different descriptive frameworks—and that a flexible, pluralistic infrastructure is more equitable than a one-size-fits-all system, however comprehensive.

For libraries managing multiple vocabularies alongside LCSH, the practical challenge is interoperability: ensuring that a patron searching a regional consortium catalog can retrieve results described using different terminological systems. This is precisely the problem that linked data is designed to solve, by establishing machine-readable equivalences between terms from different vocabularies.

What Librarians on the Ground Are Saying

Librarians navigating these transitions describe a landscape that is exciting but demanding. Cataloging units are being asked to develop fluency in linked data concepts, evaluate AI tools critically, and engage with community partners around descriptive practice—all while maintaining the day-to-day work of processing new acquisitions and supporting researchers.

Training is a significant concern. LCSH expertise is built over years of practice, and while linked data concepts are increasingly incorporated into library school curricula, many working catalogers have had to pursue professional development independently. Organizations such as the American Library Association, the Program for Cooperative Cataloging, and LD4, a community of practice focused on linked data in libraries, have expanded their training offerings, but demand continues to outpace supply.

Funding is another persistent obstacle. Migrating legacy catalog data to linked data formats, implementing AI-assisted workflows, and building partnerships with community vocabulary projects all require investment that many institutions—particularly public libraries and small academic libraries—struggle to secure.

Looking Ahead

LCSH is not disappearing. Its scale, its institutional backing, and its deep integration into catalog infrastructure ensure that it will remain central to subject access in American libraries for the foreseeable future. But the professional consensus that it should be the sole or primary framework for knowledge organization is eroding, replaced by a more pluralistic vision in which LCSH is one node in a larger, more flexible network of descriptive resources.

For researchers and library users, the practical takeaway is straightforward: the catalog you use today is almost certainly in transition, and the tools available for discovery are becoming more powerful and more varied. Librarians remain the most reliable guides through this evolving landscape—professionals who understand not only what the catalog contains but how it was built, where its limitations lie, and what alternatives exist. In an era when the infrastructure of knowledge organization is being rebuilt from the ground up, that expertise is more valuable than ever.

All Articles

Related Articles

Cataloging Culture: The Ideological Footprint Embedded in Library of Congress Subject Headings

Cataloging Culture: The Ideological Footprint Embedded in Library of Congress Subject Headings