Versions Compared

Key

  • This line was added.
  • This line was removed.
  • Formatting was changed.

...

This exercise ably demonstrates some of the limits of topic modeling. Especially in a large corpus, the topics do not necessarily encompass themes bridging various works, but instead are often just the predominant theme of one work or two works within the greater corpus. That is to say, what exactly entails a "topic" on the document or corpus level, is a variable that depends greatly on the corpus itself and requires some analysis to unveil. Another caveat would be that the "top ranked documents" within a topic (and ever perhaps the topic itself) can be skewed based on document length. To use another example from my Tractarian corpus, topic 16: "church scripture doctrine truth system fathers divine doctrines words catholic" has the most occurrences (16,970) in the 5th volume of the Tracts for the Times. But, if you zoom down another level into the breakdown of different topics within that particular volume, you will find that that particular topic constitutes only 14% of the words within the document. Out of  approximately 313,000 words, 14% equals 43,820 words. Juxtapose that against another volume titled "ElucidationsNewman," a shorter pamphlet by Newman written to address the perceived theological heterodoxy of one of his rivals. , 35% of the words in "Elucidations" are within topic 16, out of approximately 16,500 words, 35% equals 5,775 words. "Elucidations" though much shorter than the Tracts for the Times, concerns itself much more intimately with ideas of church, scripture, and doctrine than does the lengthy volume of Tracts, both in close reading and in the algorithmically generated percentages. While it may be obvious to say, ordering the importance of documents by word counts within the topics, as opposed to percentages of each topic within a document, automatically favors longer texts.

...