*** This version of Confluence is for testing only and contains a copy of content from June 29th 2026. No changes will be preserved. ***
...
The class is a guided exploration in the HTRC portal. We will explore the algorithms and use them to discover the capabilities of the algorithms, their limitations, and various strategies for addressing those challenges. This allows us to explore the HTRC portal in specific, and grapple first hand with basic issues encountered in computational analysis of text.
Before we begin
- Log on to the HTRC Production Portal with your personal credentials.
- Once signed in, click on "Algorithms" in the black navigational bar at the top of the page. Once you have done so, we will be ready to begin.
- Book mark this page for handy access. We will be referring to it at points in the workshop.
- When referencing custom stop word lists - go here: https://confluence.cornell.edu/x/BYOxDg

Resources
- HathiTrust is:
- an international partnership of over 80 institutions.
- a digital library containing over 11 million books, 33% of which are in the public domain. All items are fully indexed, allowing for full text search within all volumes.
- a trustworthy preservation repository providing long-term stewardship, redundant robust backup, continuous monitoring, persistent identifiers for all content
- HathiTrust Research Center (HTRC) - a collaborative research center (jointly managed by Indiana University and the University of Illinois) dedicated to developing cutting-edge software tools and cyberinfrastructure that enable advanced computational access to large amounts of digital text.
- HTRC Production Portal - a web-based user experience of the HTRC. The production portal makes available all the full-text indexes of the Google-digitized deposits to HathiTrust that are in the public domain.
- HTRC User Community Wiki - home of the user support documentation, meeting notes, elist addresses and sign-up information, and FAQs.