An overview of computational analysis of text, foundations, and exploration of challenges and strategies.
Prep
- bring text to play with in Voyant
Going down the rabbit hole: anatomy of a digital book
How is a digital book made? How does the structure relate to its function? What opportunities does this afford us in terms of text analysis?
(1) Consider this book: Alice's Adventures in Wonderland. Explore the controls on the right side of the page turner.
- Note the different views of the book. What types of digital files are likely to make up the parts of a digital book?
- How are these files likely to be made?
- For what types of uses are each suited?
(2) Now explore the controls at he top of the page turner.
- There is a box labeled "search in this text". What can you deduce about the book from this functionality?
- What do the other controls do? Is there a way to summarize this class of controls? What underlying logic might you predict that coordinates these functions? (Food for thought...)
(3) What might this page be? (It also has this view.) Is this also part of the book? When and how might it be used?
(4) Diving deeper into text. OCR processes are not perfect. Consider some areas of special challenge:
- poor image quality (resolution, warp, skew, crop)
- books where characters vary from "standard"
- handwritten manuscripts, block printed books
- early font styles, many non-Roman alphabets (Panjabi example)
- character-based writing systems: Japanese done well, and not so well
- language specific challenges - Arapaho gospel of St. Luke
- Mixed languages: Latin/Greek, German/Greek
- Unexpected arrangement
Computational analysis of text
We count tokens - What is tokenization? Why tokenize? What are some strategies used to tokenize?
Let's look again at the Arapaho gospel of St. Luke. Switch to text view.
- What is a word?
- What isn't a word?
- Can you think of special cases where tokens might contain more than one word?
- Are there words we prefer not to count at all?
- What sets of rules would we need to tokenize? Would these be ordered in any specific way?
- Is there a "right way" to tokenize?
What are the opportunity points that the structure and arrangement of a book afford?
- How do challenges with OCR intersect with strategies for computational analysis of text? What might be effective strategies to deal with these challenges?
- What exactly is the "text"? Can you think of parts of a book that you might not want to include in your analysis? Why or why not?
Introducing Control - "Microanalysis" and Voyant
- Voyant
- Load your text sample
- Playing with control
- observe counts
- stopwords
- What does exerting control do to our results? Does it change the validity of our assertions?
- What is signal? What is noise?
We calculate frequency
- Why not express our counts simply (as counts)? Why calculate frequencies?
- Is this misleading? If so, in what ways?
Moving from Microanalysis to Macroanalysis
- Google nGrams; help (examples)
- Bookworm; help, (examples)
- problems of convergence/divergence and strategies for disentangling.
More macroanalysis
- HTRC
- entity extraction
- What are entities?
- False positives/false negatives (omissions)
- Clustering - topic modeling
- define and defer for Mimno
Image analysis
- Edge analysis
- Neural networks
Biblio
Writing about your results - the structure of scientific papers
Matt Jockers Book