You are viewing an old version of this page. View the current version.

Compare with Current View Page History

« Previous Version 11 Next »

This page is a companion to a guest lecture prepared for ARCH 3819/ARCH5819

Class Preparation

Please bring a laptop. If you do not own one, feel free to check one out at the Olin circulation desk. Our explorations are intended to be participatory.
No special software will be needed. All explorations will be done through a Web browser.

Discussion of terms

There will be a short discussion where we define terms (at least loosely) so that we might use a common language for our discussion and critique of these tools and their strategies.

Explorations

All analysis tools will be demonstrated; no prior knowledge of any tools will be required. The room is equipped with video input to allow sharing of your desktop on the screen.  If you discover something that you would like to share and discuss, we can easily do so.

Word Frequencies

Voyant is a low barrier text analysis tool that delivers a rich, interactive interface and a variety of visualizations.

  • Data: Text of your choosing.  Upload interface accepts files in the following formats: plain text, PDF that has embedded OCR, MS Word doc and docx files, URLs. Can take multiple documents to make a set.  Upload of any material will be subject to the Voyant privacy policy.
  • Analysis: Voyant environment lists every word in the document and counts for each.  Also calculates frequencies based on the total word word count of the document(s) in the uploaded set. 
  • Visualization: Interactive dashboard: Wordcloud, graph of frequency over document segments, tabled secondary data, source data as text.  Also has navigational aids that integrate source and secondary data, and utilities like stopword control, URLs for display, etc. 

Sample texts and URLs for analysis are listed below for experimentation, but feel free to use other source data that interests you. 

nGrams

nGrams depict the frequency of a word or word phrase and are most often depicted over publication year.  We have two nGram tools, each leveraging different source data.

Google's nGram Viewer. The links below as starting points.  Dynamic modifications can be made at any point. Rules for syntax can be found on the About page.

 

Network Analysis

Data:

Analysis:

Visualization:

 

Immersion is a tool for discovering the connections in a corpus of email.  It analyzes the flow data (information found in email headers) and represents these as a network of entities.  The analysis is done in real time on the flow data for which you provide credentials.  The display is rich and  interactive. 
By design, Immersion collects only header information (From, To, Cc and Timestamp).  However, using the actual flow data from your account may cause concerns regarding privacy - Be sure to read over the FAQs to understand what information you are granting access to, and how it will be used.  If you do not like the terms of the tool, you can experience it with their demo data. 

Spatial and Temporal Representation

Data:

Analysis:

Visualization:

If time permits, we might look at Viewshare, a free tool provided by the Library of Congress that allows you to generate and customize interactive maps, timelines, facets, and tag cloud visualizations of digital collections.  The tool presupposes that you have an account and your data ready, including location and/or time related data fields in some basic forms.  A few helpful tutorials are available. Viewshare can be embedded into other web experiences. 

 

  • No labels