Text Analysis in the Wild: Steve Jobs’s Android Obsession Analyzed

I came across this example of text analysis in the wild using a wordle, Steve Jobs’s Android Obsession Analyzed. The short article is by David Zax in Fast Company (October 19, 2010.) Based on “Android” coming up as the highest frequency content word Zax reads obsession.

So yes, the Android weighs heavily upon Jobs’s mind; and his dreams are more than likely populated with ravenous green robots consuming everything in their path.

IMS Open Corpus Workbench

John pointed me to an interesting open source project, the IMS Open Corpus Workbench. This project has developed tools are for “managing and querying large text corpora (ranging from 10 million to 2 billion words) with linguistic annotations.” Obviously it has a linguistics bent, but the tools seem to be well documented and usable.

You can see an example of an interesting interface to the Corpus Workbench at BwanaNet – a wizard-like interface where you go through 5 steps to get results on an English, Catalan, and Spanish corpus.

AlchemyAPI – Transforming Text Into Knowledge

Stéfan pointed me to the AlchemyAPI service. AlchemyAPI provides an API for extracting “information about people, places, companies, topics, languages” and concepts. They have a nice demo on the front page where they take a news a top news story, extract the entities and then create a spring-loaded graph of the named entities.

You can see that for this story the system found organizations, a city, countries and persons.

A free API key is available for up to 30,000 calls a day.

Sam Winston : Darwin

I came across an artist, Sam Winston, whose work often explores language. For example Darwin (see image above) compares Darwin’s On the Origin of Species and Ruth Padel’s Darwin, A Life In Poems.

Some of the panels/pages in Darwin are visualizations, even if hand drawn.

Many of his other works also play with language and language artifacts like Folded Dictionary.

Towards a Methods Commons

Well my vacation is over and I’m facilitating a retreat on text methods across disciplines. (See Towards a Methods Commons.) With support from the ITST program at SSHRC we brought together 15 linguists, philosophers, historians, and literary scholars to discuss methods in a structured way. The goal is to sketch a commons that gathers “recipes” that show people how to do research things with electronic texts. Stay tuned for a draft web site in about 6 months.

Google: Our commitment to the digital humanities

Google has announced the first projects they are funding to use Google Books and have announced a commitment to the digital humanities of nearly a million dollars. See Official Google Blog: Our commitment to the digital humanities.

we’d like to see the field blossom and take advantage of resources such as Google Books that are becoming increasingly available. We’re pleased to announce that Google has committed nearly a million dollars to support digital humanities research over the next two years.

Society for Digital Humanities Papers

With my graduate students and colleagues I was involved in a number of papers at the SDH-SEMI The Society for Digital Humanities / La Société pour l’Étude des Médias Interactifs conference at Congress 2010 in Montreal. They included:

  • “Exclusionary Practices: A Historical Look at Public Representations of Computers in the 1950s and Early 1960s” presented by Sophia Hoosien
  • “Before the Moments of Beginning” presented by Victoria Smith
  • I presented on “Cyberinfrastructure for Research in the Humanities: Expectations and Capacity”
  • Text Analysis for me Too: An embeddable text analysis widget” presented by Peter Organisciak
  • Daniel Sondheim talked about the interface of the citation from print to the web as part of a panel on INKE Interface Design.
  • “Theorizing Analytics” was presented by Stéfan Sinclair
  • “Academic Capacity in Canada’s Digital Humanities Community: Opportunities and Challenges” was presented by Lynne Siemens
  • “What do we say about ourselves? An analysis of the Day
    of DH 2009 data” was presented by Peter Organisciak
  • and I presented on “The Unreality of the Timeline” as part of a panel on temporal modeling at the CHA

As the papers get posted, I’ll blog them.

U of A text mining project could help businesses

Well, I made it into the computer press in Canada. An article on the Digging Into Data project I am working on has been published, see U of A text mining project could help businesses (Rafael Ruffolo, March 25, 2010 for ComputerWorld Canada.)

It is always interesting to see what the media find interesting in a story. They usually have a better idea of what their audience wants to read about so they adapt for that audience.

Who’s your DH Blog Mate: Match-Making the Day of DH Bloggers with Topic Modeling

On the 18th of March we ran the second Day of Digital Humanities, which seems to have been a success. We had more participants and some interesting analysis. Matt Jockers, for example, tried Latent Dirichlet Allocation on the blogs and wrote up the results on his blog in a post,  Who’s your DH Blog Mate: Match-Making the Day of DH Bloggers with Topic Modeling. Neat!

The General Inquirer

Reading John B. Smith’s “Computer Criticism”, (Style: Vol. XII, No. 4) I came a reference to a content analysis program called the The General Inquirer from the 1960s. This program still has a following and has been rewritten in Java. See the Inquirer Home Page. There is a web version where you can try it here (DO NOT USE A LARGE TEXT).

The General Inquirer “maps” a text to a thesaurus of categories, disambiguating on the way. The web page about How the General Inquirer is used describes what it does thus:

The General Inquirer is basically a mapping tool. It maps each text file with counts on dictionary-supplied categories. The currently distributed version combines the “Harvard IV-4” dictionary content-analysis categories, the “Lasswell” dictionary content-analysis categories, and five categories based on the social cognition work of Semin and Fiedler, making for 182 categories in all. Each category is a list of words and word senses. A category such as “self references” may contain only a dozen entries, mostly pronouns. Currently, the category “negative” is our largest with 2291 entries. Users can also add additional categories of any size.

As they say later on, their categories were developed for “social-science content-analysis research applications” and not for other uses like literary study. The original developer published a book on the tool in 1966:

Philip J. Stone, The General Inquirer: A Computer Approach to Content Analysis. (Cambridge: M. I. T. Press, 1966).