What are we doing? A Deeper Dive into TEI and Text Analysis

Text analysis is a broad term that encompasses the examination and interpretation of textual data. It involves various techniques to understand, organize, and derive insights from text, including methods from linguistics, statistics, and machine learning. Text analysis often includes processes like text categorization, sentiment analysis, and entity recognition, to gain valuable insights from textual data. (Definition from George Washington University Library)

It’s not uncommon for anyone who hears about The Digital Gogol Project to be confused about what it is we’re actually doing. Explaining that we’re encoding a text using XML markup language so that a computer can “read” it and that we can then query it and visualize it as data is not quite sufficient. The terminology, coming from computer science, is unfamiliar to humanities scholars. There are concerns about the intrusion of computers into humanities work. A machine can’t close-read a story, can’t understand emotion, motivation, and narrative play, or interpret intention, nuance or context – in short, all the elements of what it means to be human in this world. Computers are good with data, and most people might say that literature and history are the opposite of data, so what are computers doing in humanities work? 

To find the answer, we must first turn our texts into data.

Within the digital humanities (DH) context, text encoding is a long-established tradition coming out of the possibilities offered by emerging digital technologies. However, as libraries, museums, and individual scholars began looking for ways to represent texts in digital form, new platforms, systems, and standards proliferated. The problem was that these systems were competing, unstandardized, and often inadequate, impeding the ability of GLAM institutions and scholars to share materials in a long-standing and sustainable way. The idea for the Text Encoding Initiative (TEI) was developed first in 1987 during a scholarly “meeting of the minds” sponsored by the Association for Computers in the Humanities, and the first guidelines were produced in 1990, with subsequent revisions and updates occurring periodically ever since. TEI guidelines are maintained and promoted by the TEI Consortium and its standards are widely recognized as an important tool for preserving textual information in digital form.

Initially, the TEI operated in the realm of scholarly digital editions and archives. Scholars would encode manuscripts, rare texts, and cultural heritage materials and make them available digitally as searchable documents. Accessibility has always been at the forefront of this initiative and the TEI has helped give many lesser known texts and archival collections a new digital life. Similarly, I wanted to give Gogol’s work a new digital life by providing XML files of his works so that other scholars and researchers can use them for their own digital humanities work. XML (eXtensible Markup Language) is just an easy and flexible way of storing data. It organizes and describes information with tags and allows for efficient data sharing. Its standardized structure eliminates the need to “translate” data between applications with custom code, because the data structure is already contained in the data. The Text Encoding Initiative is a version of XML that is tailored to literary and historical texts with specific tags relevant to those sources. 

By making the files of Gogol’s works in the original Russian available publicly as XML files, anyone can download and use them in their projects, bypassing the (sometimes) tedious process of encoding the texts from scratch. Beyond creating these files, once the texts of Gogol’s stories have been turned into data, we will then visualize that data to find patterns, frequencies, trends, networks, connections between texts. We will zoom out to make visible some details that we cannot see with close reading; we can analyze hundreds of texts simultaneously. We can also use any of a number of digital humanities tools and platforms in combination with programming languages to perform sentiment analysis, topic modelling, and mapping. 

My hope in doing this is to advance the field of Slavic digital humanities. Slavic studies has been slower to adopt the digital approach than other fields, even though new technologies offer innovative ways of engaging with these materials. There are misconceptions about what DH is and suspicions about the role of the digital in humanities work, not to mention the lack of institutional support and the learning curve required to implement an infrastructure that supports this kind of work. The digital humanities do not replace traditional humanities skills and approaches, but a complement to them. The human component cannot be eliminated, because any computational results still require a human for interpretation and contextualization. 

What about AI?

It’s an inescapable question. I do not like artificial intelligence and try to avoid using it, even though I acknowledge that there are situations and problems that could benefit from AI usage. The digital humanities is not the same as artificial intelligence, though there are areas of overlap as well as a growing subfield of digital humanities scholars who make active use of it. Some digital humanists enlist the help of artificial intelligence tools to accomplish the bulk of the basic structural tagging, and I see why. I do not plan to enlist artificial intelligence to accomplish any part of this project, no matter how tedious the work gets. I’m not looking for quick results. Although many of the computational methods of the digital humanities fall under the loose umbrella term of “distant reading,” encoding one of Gogol’s stories is about working closely with the text. When I am deliberating about which elements to tag and how to tag them, I’m evaluating the text, catching new details and nuances, and posing new questions about my own interpretation. 

The Digital Gogol Project is one of textual inquiry and possibility.  At a time when AI is available with easy ready-made answers, we’re creating a space for investigation, for making mistakes, for following whims and going down rabbit holes. We’re learning new skills and testing theories and allowing curiosity to determine our focus and next steps. There is no true end result, only experimentation with different tools and perspectives. I’m not sure what we’ll find, but I’m sure it will be worthwhile!