Versions Compared

Key

  • This line was added.
  • This line was removed.
  • Formatting was changed.

...

To further support this assumption, we can zoom in yet one more level within the Mallet results. Clicking on the link to "Lectures on Justification" yields a sample of the text itself along with a list of the most frequent topics occurring within the volume. Unsurprisingly, we find that topic 1 is far and away the most dominant topic, as 48% of the words within the document were assigned to topic 1 whereas the next closest, topic 20 from our main index "god christ world man lord holy men day faith st" is a distant second at 12%

 

This exercise ably demonstrates some of the limits of topic modeling. Especially in a large corpus, the topics do not necessarily encompass themes bridging various works, but instead are often just the predominant theme of one work or two works within the greater corpus. That is to say, what exactly entails a "topic" on the document or corpus level, is a variable that depends greatly on the corpus itself and requires some analysis to unveil. Another caveat would be that the "top ranked documents" within a topic (and ever perhaps the topic itself) can be skewed based on document length. To use another example from my Tractarian corpus, topic 16: "church scripture doctrine truth system fathers divine doctrines words catholic" has the most occurrences (16,970) in the 5th volume of the Tracts for the Times. But, if you zoom down another level into the breakdown of different topics within that particular volume, you will find that that particular topic constitutes only 14% of the words within the document. Out of  approximately 313,000 words, 14% equals 43,820 words. Juxtapose that against another volume titled "ElucidationsNewman," a shorter pamphlet by Newman written to address the perceived theological heterodoxy of one of his rivals, 35% of the words in "Elucidations" are within topic 16, out of approximately 16,500 words, 35% equals 5,775 words. "Elucidations" though much shorter than the Tracts for the Times, concerns itself much more intimately with ideas of church, scripture, and doctrine than does the lengthy volume of Tracts, both in close reading and in the algorithmically generated percentages. While it may be obvious to say, ordering the importance of documents by word counts within the topics, as opposed to percentages of each topic within a document, automatically favors longer texts.

...

One of the drawbacks to the Mallet GUI system is that it does not tell you about relationships between topics. Especially in a larger corpus with many documents and many topics, it's very difficult to go through by hand and figure out which topics seem to co-occur. Ability to look at co-occurrence of topics throughout the corpus would provide yet another layer into the idea of thematized searching that Topic Modeling embodies. As a corollary to that, it would be helpful if you could read co-occuring topics across documents as well. Of course, you can see on the "top ranked documents" page the documents and how frequently one particular topic occurs in each of them, but it would be interesting to see in which documents two topics tend to co-occur.

Additionally, while Mallet quickly and ably produces topics for one volume of work (it took two seconds to work through Darwin's On the Origin of Species), the rest of its functionality as an analytic tool suffers. For instance, running the program on Origin of Species gives a standard list of 20 topics ranging from "life conditions change ancient animals sterility struggle existence instincts parents" to "plants birds insects productions numbers early area seeds tendency perfectly" but when you click on the topic to see where it appears in the documents you get a dizzying list of over 500 different locations. Clicking on one of the "documents," leads you to a text fragment within the volume. Taken out of context, these text fragments are generally too small to be of much help.

 

Image Added

Image Added

For those curious, my results are available online to play with, although it is an ongoing experiment so they may or may not always be up: http://digitalmedia.library.cornell.edu/digital_humanities/output_html/all_topics.html, happy topic modeling!