Literary Analysis and the Wolfram Language

digital-research-methods-cover-2015-medium

Lately I’ve been trying Wolfram Mathematica more an more for analytics. I was introduced to Mathematica by Bill Turkel and Ian Graham who have done some impressive stuff with it. Bill Turkel has now created a open access, open content, and open source textbook Digital Research Methods with Mathematica. The text is a Mathematica notebook itself so, if you have Mathematica you can actually use the text to do analytics on the spot.

Wolfram has also posted an interesting blog entry on Literary Analysis and the Wolfram Language: Jumping Down a Reading Rabbit Hole. They show how you can generate word clouds and sentiment analysis graphs easily.

While I am still learning Mathematica, some of the features that make it attractive include:

  • It uses a “literate programming” model where you write notebooks meant to be read by humans with embedded code rather than writing code with awkward comments embedded.
  • It has a lot of convenient Web, Language, and Visualization functions that let you do things we want to do in the digital humanities.
  • You can call on Wolfram Alpha in a notebook to get real world knowledge like capital cities or maps or language information.

philosophi.ca : Digital Humanities Concepts 2015

TU Darmstadt MA LLC Structure

Just left a most delightful conference on Key ideas and concepts of Digital Humanities in Darmstadt, Germany. My conference notes are on philosophi.ca : Digital Humanities Concepts 2015. The conference brought together an extraordinary set of speakers who were influential in the field when I entered it. Susan Hockey, Michael Sperberg-McQueen, Nancy Ide, George Landow, Wilhelm Ott and the list goes on. I would be hard pressed to imagine a conference I have been at better able to reflect on the history and ideas of humanities computing. The organizers Andrea Rapp, Michael Sperberg-McQueen, Sabine Bartsch and Michael Bender deserve much more praise than I was able to lavish on them.

Among all the great papers I will mention:

  • Michael Sperberg-McQueen gave a very smart and well argued paper on descriptive markup arguing against its dismissal as enforcing hierarchies.
  • Marco Passarotti talked about the Index Thomisticus (which he directs) and the Busa Archive. He brought some documents including some Gantt charts and early letters. I am definitely going to visit him and the archive in Milan.
  • Fotis Jannidis gave a great paper on topic modelling and its temptations. He has very interesting stuff to say about how the method has been adopted by humanists.
  • Julia Flanders gave a paper on “Looking for Gender in the History of DH” that when published will, I predict, become mandatory reading. She gives us a way forward after what happened at DH 2015. It was a truly wise and humble talk that could go a long way to providing an inclusive way forward.
  • Nancy Ide gave a great overview of the separate trajectories taken by DH and Corpus Linguistics.
  • Peter Robinson gave a call for open editions and walked us through what that might mean.

Given the speakers, there was a lot of reflection on the history of humanities computing and disciplinarity, though enframed by a German context. TU Darmstadt has an MA in Linguistic and Literary Computing (see image of the structure of the degree above) and is now developing an undergrad degree.

Text Mining The Novel 2015

novelTMworkshop

On Thursday and Friday (Oct. 22nd and 23rd) I was at the 2nd workshop for the Text Mining the Novel project. My conference notes are here Text Mining The Novel 2015. We had a number of great papers on the issue of genre (this year’s topic.) Here are some general reflections:

  • The obvious weakness of text mining is that it operates on the novel as text, specifically digital text (or string.) We need to find ways to also study the novel as material object (thing), as a social object, as a performance (of the reader), and as an economic object in a market place. Then we also have to find ways to connect these.
  • So many analytical and mining processes depend on bags of words from dictionaries to topics. Is this a problem or a limitation? Can we try to abstract characters, plot, or argument.
  • I was interested in the philosophical discussions around the epistemological in novels and philosophical claims about language and literature.

 

DH 2015 in Sydney, Australia

Digital Humanities 2015 (DH2015) is now finishing up. I have been keeping my conference notes here.

The conference was held on the lovely campus of the University of Western Sydney. I was part of a couple of events and papers at this conference including:

  • News Scholars Symposium: With Rachel Hendry, I helped organize a pre-conference event for new scholars. This was supported by CHCI, centerNet, the Kule Institute for Advanced Study and the University of Western Sydney.
  • I participated in a public panel on Building Communities and Networks in the Humanities where I talked about some of the forms of public engagement that we are trying at the Kule Institute including the Around the World Conference.
  • I helped Stéfan Sinclair with a workshop on Voyant 2.0 (link goes to current version which will soon be 2.0).
  • I gave a paper with Stéfan Sinclair on “Talking about Programming the Digital Humanities” that traced a history of the discussion about programming and tools in the digital humanities.
  • Finally, John Montague gave a paper on “Exploring Large Datasets with Topic Model Visualizations” that I was involved in. This paper discussed a visualization for exploring the results of topic modelling that you can try in prototype here.

It is hard to summarize a whole conference, but I would note some of the questions that the new scholars posed in the unconference are worth thinking about:

  • How does one learn about the field of digital humanities?
  • How does one learn skills in the digital humanities?
  • How does one teach the digital humanities?
  • What are the ethical issues in digital work in the humanities?

Editorialisation Et Nouvelles Formes De Publication

In the last couple of weeks I’ve been at two interesting conferences and took research notes.

  1. I gave a keynote on “Big Data and the Humanities” at the Northwestern Research Computation Day (link to my research notes). I gave a lot of examples of projects and visualizations.
  2. At the Éditorialisation Et Nouvelles Formes De Publication (link to my research notes) conference I spoke about “Publishing Tools: A Theatre of Machines”. I showed how text analysis machines have evolved.

Is it Research or is it Spying? Thinking-Through Ethics in Big Data AI and Other Knowledge Sciences

Is it Research or is it Spying? Thinking-Through Ethics in Big Data AI and Other Knowledge Sciences has just been published online. It was written with Bettina Berendt and Marco Büchler and came out of a Dagschule retreat where a group of us started talking about ethics and big data. Here is the abstract:

How to be a knowledge scientist after the Snowden revelations?” is a question we all have to ask as it becomes clear that our work and our students could be involved in the building of an unprecedented surveillance society. In this essay, we argue that this affects all the knowledge sciences such as AI, computational linguistics and the digital humanities. Asking the question calls for dialogue within and across the disciplines. In this article, we will position ourselves with respect to typical stances towards the relationship between (computer) technology and its uses in a surveillance society, and we will look at what we can learn from other fields. We will propose ways of addressing the question in teaching and in research, and conclude with a call to action.

A PDF of our author version is here.

Wilkens: Literary Attention Lag

Matthew Wilkens has posted a nice blog essay about his short MLA paper on geography and memory, Literary Attention Lag. He looked at how some cities get far more literary attention than their population merits despite a general correlation between population and attention. For example, in 1860 Chicago and New Orleans had about the same population, but New Orleans gets a lot more attention.

What is particularly useful is that he provides an iPython notebook with a documented version of his code here. He also provides a link to his data so you can edit and recapitulate his study.

Stéfan Sinclair and I are experimenting with Mathematica notebooks and iPython notebooks as a way to share research thinking with code woven in.

Is GamerGate About Media Ethics or Harassing Women? Harassment, the Data Shows

PeopleTargeted

In all the GamerGate stories, an interesting move by Newsweek as to commission a study of GamerGate tweets. Taylor Wofford reported about the results in an article from October 25th, 2014 that is titled, Is GamerGate About Media Ethics or Harassing Women? Harassment, the Data Shows. The study was run by BrandWatch  and they looked at who was the target of tweets with #gamegate. Low and behold the GamerGate community seemed more concerned with female game designers than journalists which calls into question the claim that GamerGate is really about ethics and games journalism.

We are now gathering tweets too and we will see if we can reproduce the results. At first glance the number of GamerGate tweets seems really low – they seem to be sampling. It will also be interesting to see if there has been a shift in emphasis in the discussion.

bookworm

chart (1)The folks behind the Google Ngram Viewer have developed a new tools called bookworm. It has a number of corpora (the example above is from bills from beta.congress.gov.) It lets you describe more complex queries and you can upload your own data.

Bookworm is hosted by the Cultural Observatory at Harvard directed by Erez Lieberman Aiden and Jean-Baptiste Michel who were behind the NGgam Viewer. They have recently published a book Uncharted where they talk about different cultural trends they studied using the NGram Viewer. The book is accessible though a bit light.

Checktext.org

I was sent a note about Checktext.org, an web site where you paste (or upload) some text and it gives you basic analytical information like Flesch-Kincaid Grade Level. One neat feature is that it will do a plagiarism check against a database. It isn’t clear how they build their database or if they are just using Google, but it caught a web page I passed it.