Searching Across Texts: Introducing the Han-Nom Corpus Search in Han-Nom Research Hub

September 22, 2026
Searching Across Texts: Introducing the Han-Nom Corpus Search in Han-Nom Research Hub

For centuries, scholars working with Hán-Nôm materials have approached texts through close reading: locating a work, turning through its pages, identifying a passage, and following words, names, and concepts from one source to another. Digitization has made many of these materials accessible online, but access to digital images is only the beginning. What becomes possible when we can search across texts as well as read them individually?

 

The Corpus Search function in Digitizing Việt Nam’s Hán-Nôm Research Hub is designed to make this possible.

 

Part of the broader Hán-Nôm Research Hub, Corpus Search provides a full-text search environment for exploring digitized Hán-Nôm materials. Instead of beginning only with the title of a known work, users can begin with a word, character, name, phrase, or concept and search for its occurrences across the texts currently available in the corpus. 

 

From a digital library to a research corpus

 

This represents an important shift in how a digital collection can be used.

 

A conventional digital library is particularly effective when we already know what we are looking for. A researcher may search for a particular title, author, historical period, or manuscript and then open the corresponding digital object. The Hán-Nôm collections accessible through Digitizing Việt Nam already make a substantial body of material discoverable in this way, bringing together digitized Hán-Nôm texts and other materials originating from the Vietnamese Nôm Preservation Foundation and the National Library of Vietnam and hosted through Columbia University’s Digital Library Collections.

 

Corpus Search adds another layer. It allows the collection itself to become an object of inquiry.

 

A researcher interested in a particular term, for example, can search across available texts to identify where that term appears. A historian might trace references to a person, place, institution, or administrative concept. A scholar of literature might investigate recurring vocabulary or expressions across different works. A researcher of intellectual history might begin exploring how particular concepts occur across authors, genres, or periods.

 

The question therefore changes from:

“Where can I find this book?”

to:

“Where does this idea, word, person, or expression appear across the collection?”

 

That seemingly small change opens a different mode of working with historical texts.

 

How Corpus Search works

 

From the Hán-Nôm Research Hub, users can enter a query directly into the Corpus Search box. The system performs full-text searching across the corpus of classical Vietnamese literature and historical documents. The dedicated Hán-Nôm Corpus Database provides additional ways of navigating the materials, including full-text keyword searching and an instant title search. Users can also browse and filter the library according to bibliographic information such as genre, language, contributor, and year.

 

Importantly, the full-text search can retrieve matching excerpts in Hán-Nôm or Quốc ngữ, allowing different forms of textual inquiry to begin from the same research environment. 

 

These functions complement one another. A scholar may begin with a broad keyword search, notice that an expression appears in several works, narrow the materials by period or other metadata, and then return to individual texts for closer examination.

Corpus Search is therefore not intended to replace close reading. Rather, it can help researchers decide where to read closely.

 

Seeing connections that are difficult to see one text at a time

 

The scholarly potential of corpus search becomes especially apparent when working with large collections.

 

Imagine, for example, a researcher interested in the concept of “heavenly mandate” (天命). Traditionally, investigating the term across a large body of Hán-Nôm sources could require identifying potentially relevant works and examining them individually. Corpus Search offers another starting point: search for the term across the available corpus, examine the passages in which it occurs, identify the texts containing those passages, and then investigate those texts in their historical and textual contexts.

 

The same approach could be applied to place names, official titles, religious vocabulary, literary expressions, personal names, medical terminology, or concepts such as 忠 (loyalty), 孝 (filial piety), 國 (country/state), or 文 (writing/culture).

 

For experienced Hán-Nôm scholars, this can accelerate the exploratory stage of research and reveal connections worth investigating further. For students and researchers entering the field, it offers another point of entry into collections whose scale, scripts, and bibliographic traditions can otherwise make exploration difficult.

 

Search is a beginning, not an interpretation

 

Corpus searching also requires scholarly caution.

 

Finding the same sequence of characters in several texts does not necessarily mean that they carry precisely the same meaning. Historical vocabulary changes over time; characters can have multiple readings; textual traditions contain variants; and the significance of a term depends on its surrounding passage, genre, author, and historical circumstances. Digitized and machine-readable historical texts may also contain transcription or recognition errors.

 

For these reasons, search results should be understood as research leads rather than conclusions.

 

The value of Corpus Search lies precisely in this relationship between computational discovery and humanistic interpretation. Digital tools can help scholars locate patterns and connections across a scale of material that would be difficult to examine manually. It remains the task of researchers to return to the sources, read passages in context, compare textual witnesses where necessary, and determine what those connections actually mean.

 

Building new ways into Vietnam’s textual heritage

 

Corpus Search is part of a larger experiment behind Digitizing Việt Nam’s Hán-Nôm Research Hub. The Hub brings digital archives together with tools including automated Hán-Nôm OCR, unified dictionary lookup, manuscript and textual research resources, and learning tools. Its purpose is not simply to place historical materials online, but to create different pathways through which those materials can be discovered, studied, taught, and reused. 

 

For the broader public, this can make an unfamiliar textual heritage more approachable. For students, it creates opportunities to learn by exploring actual historical sources. For scholars, it provides a bridge between established practices of philology and close reading and emerging methods of corpus-based and computational research.

 

Most importantly, Corpus Search invites users to ask new questions of old texts.

 

A digital archive preserves a document and makes it accessible. A searchable corpus goes one step further: it allows us to begin seeing relationships between documents—between words, people, places, concepts, and texts that may have been produced decades or centuries apart.

 

The results do not provide the final interpretation. They provide somewhere to begin.

 

Explore the Hán-Nôm Research Hub and try Corpus Search on the Digitizing Vietnam Platform.