Voyant at Georgia Tech

Today I Skyped into a class by Lauren Klein on Digital Humanities at Georgia Tech. The students all had to use Voyant for an assignment and they had a great set of questions to ask me. See Questions for Professor Rockwell.

Klein also had her students post short essays on using Voyant on Sherlock Holmes under the category Sherlock Holmes Text Analysis. You can see the range of reactions from frustration with the tool, to “so what”, to students who find the “surfing and stumbling” creative. I’m impressed at how Professor Klein has put together a reasonable exercise in text analysis for undergrads.

In the spirit of Voyant, here is a word cloud of the student assignments on the course blog:


WordSeer

Stéfan pointed me to Berkley WordSeer a text analysis tool “that includes visualizations and works on the grammatical structure of text.” You can watch the video with Aditi Muralidharan talking about the project. She sees the problem with traditional search being the way keyword reading models texts as a bag of words. What we can’t do is model text as sentences. In other words she wants to leverage natural language processing to enhance search so you can see how “God” is described or what she/he has done. There are also some visualization tools like a heat map and word tree.

There is a nice YouTube video demoing how to use WordSeer to explore “beautiful” in Shakespeare.

Antconc – Concordance tool on PC/Mac

Screen shot of Antconc

Thanks to John, I learned about a gem of a concordance tool for the Mac, PC and Linux called Antconc. It runs on your computer and you can download the tool from the author’s site, Laurence Anthony’s Software. If it is stable it could be a great tool to introduce students to text analysis. Looking at the screenshots it has some nice features for finding n-grams and can handle a set of texts.

Analysis of 250,000 hacker conversations

 

From Slashdot a story about the text Analysis of 250,000 hacker conversations. A security company Imperva has been analyzing hacker forums to understand trends, how people learn about hacking, and what are popular strategies.

In the Imperva report, Hacker Intelligence Initiative, Monthly Trends Report #5 (PDF) they describe their methodology as “content analysis” (their quotations) but it mostly involves searching for threads and reading. The report has great examples of the types of discussions.

A good example of how simple text analysis can help industry understanding.

Every story has a beginning

Every story has a beginning is the text of a keynote by Tim Sheratt that nicely weaves individual stories together as an example of what we can do with information technology. I highly recommend it; he quotes Steve Ramsay and Tim Hitchcock to the effect that what is important are the stories of individuals like those he paints through the digital archives he has access to. He sets this humanistic view of how we can use the technology against the Culturomics approach which is trying to turn history and its archives into grist for cultural science. Sheratt calls the culturomic vision “barren” and I tend to agree. He ends by asking,

But who defines the problems?

His answer is Linked Data which “gives us a way to present an alternative to Google’s version of the world. We can argue back against the search engines, defining our own criteria for relevance, and building our own discovery networks.” (And his talk has a link for those who want to view the triples…) I would say that we can also build tools like Voyant (formerly Voyeur, which he uses) to help us begin to tell the stories.

Canadian Writing Research Collaboratory Launch

 

I am at the Canadian Writing Research Collaboratory (CWRC) launch. CWRC is building a collaborative editing environment that will allow editorial projects to manage the editing of electronic scholarly editions. Among other things CWRC is developing an online XML editor, a editorial workflow management tools, and integrated repository.

The keynote speakers for the event include Shawna Lemay and Aritha Van Herk.

Happy Words Trump Negativity in the English Language

Happy Words Trump Negativity in the English Language is an interesting story about a study by Kloumann and colleagues on Positivity of the English Language. They used Mechanical Turk to get people to assess whether the high frequency words used in Twitter, Books, the New York Times and Music Lyrics were positive. Their study showed that overwhelmingly English is a positive language. Thanks to Stan for this.

Old Bailey Trials Are Tabulated for Scholars Online

The New York Times now has an article on the Criminal Intent project I was part of. See, Old Bailey Trials Are Tabulated for Scholars Online. They quote a historian who is sceptical of the results of mining, though he appreciates the resource.

“The Old Bailey Online project has done a great service in making those sources widely (and costlessly) available,” Mr. Langbein wrote in an e-mail. But he complained that the claims about data mining have “a breathless quality: ‘you can expect big things from us,’ but as yet it’s all method and no results.” He said that the new findings belittle the work of a generation of scholars who focused on the 18th century as the turning point in the evolution of the criminal justice system.

Alas, he seems didn’t read our report, but the summary in the Chronicle. It is easy to use cute phrases like “breathless quality”, but is he right? Time will tell, but I think the historians on our team have backed up the results found with mining and they never belittled the work of previous scholars – we saw ourselves building on it.

What can mining do? I think mining can give you a big picture so that you see the forest rather than trees in a way that no one could before. Conclusions about the shape of the forest have to be checked against other evidence, but the results of mining is evidence that is not breathless even if it takes your breath away. As Bill Turkel put it,

Mr. Turkel, who developed some of the digital tools, said that data mining reveals unexpected trends and connections that no one would have thought to look for before. Previous scholars “tended to cherry-pick anecdotes without having a sense that it was possible to measure all of that text and treat the whole archive as a single unit,” he said.

Of course, if you then leverage traditional evidence to buttress your argument then the mining is forgotten or trivialized.