Brad, a student who has been working on some cool text analysis stuff, pointed me to Facebook | Lexicon. In Lexicon you type in one or more words and it will graph their popularity over time in Facebook walls. I entered “theatre, games, literature, movies” – guess which two were far more popular?
Category: Text Analysis
VersionBeta3 < Main < WikiTADA
We have a new version of the Big See collocation centroid. Version Beta 3 now has a graphical user interface where you can control settings before running the animation and once the animation is run. As before we show the process of developing the 3D model as an animation. Once run you can manipulate the 3D model. If you turn on stereo you can see the text model as a 3D object if you have the right glasses on (it supports different types including red/green.)
I’m still trying to articulate the goals of the project. Like any humanities computing project the problem and solutions are emerging as we develop and debate. I now think of it as an attempt to develop a visual model of a text that can be scaled out to very high resolution displays, 3D displays, and high performance computing. The visual models we have in the humanities are primitive – the scrolling page and the distribution graph. TextArc introduced a model, the weighted centroid, that is rich and rewards exploration. I’m trying to extend that into three dimensions while weaving in the distribution graph. Think of the Big See is a barrel of distributions.
High Resolution Visualization

In a previous post I wrote about a High Performance Visualization project. We got the chance to try the visualization on a Toshiba high resolution monitor (something like 5000 X 2500). Above you can see a picture I took with my Blackberry.
What can we do with high resolution displays? What would we show and how could we interact with them? I take it for granted that we won’t just blow up existing visualizations.
High Performance Visualization
I’m working with the folks at our local HPC consortium, SHARCNET on imagining how we could visualize texts with high resolution displays, 3D displays, and cluster computing. The project, temporarily called The Big See has generated an interested beta version. You can see a video on the process running and images from the final visualization here, Version Beta 2.
One of the unanticipated insights from this project is that the process of building the 3D model, which I will call the *animation*, is as interesting as the final visual model. From the very first version you could see the text flowing up and the high frequency words jostling each other for position. Words would start high and then slide clockwise around. Collocations build up as it goes. We don’t have the animation right, but I think we are on to something. You can see Version B2 as an MP4 animation here.
Now we will start playing with the parameters – colours, transparency, and weight of lines.
Next Steps for E-Science and the Textual Humanities
D-Lib Magazine has a report on next steps for high performance computing (or as they call it in the UK, “e-science”) and the humanities, Next Steps for E-Science, the Textual Humanities and VREs. The report summarizes four presentations on what is next. Some quotes and reactions,
The crucial point they made was that digital libraries are far more than simple digital surrogates of existing conventional libraries. They are, or at least have the potential to be, complex Virtual Research Environments (VREs), but researchers and e-infrastructure providers in the humanities lack the resources to realize this full potential.
I would call this the cyberinfrastructure step, but I’m not sure it will be libraries that lead. Nor am I sure about the “virtual” in research environments. Space matters and real space is so much more high-bandwidth than the virtual. In fact, subsequent papers made something like this point about the shape of the environment to come.
Loretta Auvil form the NCSA is summarized to the effect that Software Environment for the Advancement of Scholarly Research (SEASR) is,
API-driven approach enables analyses run by text mining tools, such as NoraVis (http://www.noraproject.org/description.php) and Featurelens (http://www.cs.umd.edu/hcil/textvis/featurelens/) to be published to web services. This is critical: a VRE that is based on digital library infrastructure will have to include not just text, but software tools that allow users to analyse, retrieve (elements of) and search those texts in ever more sophisticated ways. This requires formal, documented and sharable workflows, and mirrors needs identified in the hard science communities, which are being met by initiatives such as the myExperiment project (http://www.myexperiment.org). A key priority of this project is to implement formal, yet sharable, workflows across different research domains.
While I agree, of course, on the need for tools, I’m not sure it follows that this “requires” us to be able to share workflows. Our data from TAPoR is that it is the simple environment, TAPoRware, that is being used most, not the portal, though simple tools may be a way in to VREs. I’m guessing that the idea of workflows is more of a hypothesis of what will enable the rapid development of domain specific research utilities (where a utility does a task of the domain, while a tool does something more primitive.) Workflows could turn out to be perceived of as domain-specific composite tools rather than flows just as most “primitive” tools have some flow within them. What may happen is that libraries and centres hire programmers to develop workflows for particular teams in consultation with researchers for specific resources, and this is the promise of SEASR. When it crosses the Rubicon of reality it will provide support units a powerful way to rapidly deploying sophisticated research environments. But if it is programmers who do this, will they want a flow model application development environment or default back to something familiar like Java. (What is the research on the success of visual programming environments?)
Boncheva is reported as presenting the Generic Architecture for Text Engineering (GATE).
A key theme of the workshop was the well documented need researchers have to be able to annotate the texts upon which they are working: this is crucial to the research process. The Semantic Annotation Factory Environment (SAFE) by GATE will help annotators, language engineers and curators to deal with the (often tedious) work of SA, as it adds information extraction tools and other means to the annotation environment that make at least parts of the annotation process work automatically. This is known as a ‘factory’, as it will not completely substitute the manual annotation process, but rather complement it with the work of robots that help with the information extraction.
The alternative to the tool model of what humanists need is the annotation environment. John Bradley has been pursuing a version of this with Pliny. It is premised on the view that humanists want to closely markup, annotate, and manipulate smaller collections of texts as they read. Tools have a place, but within a reading environment. GATE is doing something a little different – they are trying to semi-automate linguistic annotation, but their tools could be used in a more exploratory environment.
What I like about this report is we see the three complementary and achievable visions of the next steps in digital humanities:
- The development of cyberinfrastructure building on the library, but also digital humanities centres.
- The development of application development frameworks that can create domain-specific interfaces for research that takes advantage of large-scale resources.
- The development of reading and annotation tools that work with and enhance electronic texts.
I think there is fourth agenda item we need to consider, which is how we will enable reflection on and preservation of the work of the last 40 years. Willard McCarty has asked how we will write the history of humanities computing and I don’t think he means a list of people and dates. I think he means how we will develop from a start-up and unreflective culture to one that one that tries to understand itself in change. That means we need to start documenting and preserving what Julia Flanders has called the craft projects of the first generations which prepared the way for these large scale visions.
Toy Chest (Online or Downloadable Tools for Building Projects)
Alan Liu and others have set up a Knowledge Base for the Department of English at UCSB which includes a neat Toy Chest (Online or Downloadable Tools for Building Projects) for students. The idea is to collect free or very cheap tools students can use and they have done a nice job documenting things.
The idea of a departmental knowledge base is also a good one. I assume the idea is that this can be an informal place for public knowledge faculty, staff and students gather.
netzspannung.org | Archive | Archive Interfaces

netzspannung.org is a German new media group with an archive of “media art, projects from IT research, and lectures on media theory as well as on aesthetics and art history.” They have a number of interfaces to this archive, for an explanation see, Archive Interfaces. The most interesting is the Java Semantic Map (see picture above.)
netzspannung.org is an Internet platform for artistic production, media projects, and intermedia research. As an interface between media art, media technology and society, it functions as an information pool for artists, designers, computer scientists and cultural scientists. Headed by » Monika Fleischmann and » Wolfgang Strauss, at the » MARS Exploratory Media Lab, interdisciplinary teams of architects, artists, designers, computer scientists, art and media scientists are developing and producing tools and interfaces, artistic projects and events at the interface between art and research. All developments and productions are realised in the context of national and international projects.
See The Semantic Map Interface for more on their Java Web Start archive browser.

We feel fine: Blog harvesting
We Feel Fine is “an exploration of human emotion, in six movements” that harvests recent blog posts for “I feel” or “I am feeling” and then stores information for visualization. There is an applet where you can look for “men who are abiding”. Thanks to Guy for this.
Visuwords: Visual Dictionary Graph
Visuwords online graphical dictionary and thesaurus is a tool that visualizes WordNet relationships. It shows the synonyms and definitions for words you search for. Thanks to Shawn for this.
OpenSocial – Google Code
Two days ago, on the day of All Hallows (All Saints), Google announced OpenSocial a collection of APIs for embedded social applications. Actually much of the online documentation like the first OpenSocial API Blog entry didn’t go up until early in the morning on November 2nd after the Campfire talk. On November 1st they had their rather hokey Campfire One in one of the open spaces in the Googleplex. A sort of Halloween for older boys.

Screen from YouTube video. Note the campfire monitors.
OpenSocial, is however important to tool development in the humanities. It provides an open model for the type of energetic development we saw in the summer after the Facebook Platform was launched. If it proves rich enough, it will provide a way digital libraries and online e-text sites can open their interface to research tools developed in the community. It could allow us tool developers to create tools that can easily be added by researchers to their sites – tools that are social and can draw on remote sources of data to mashup with the local text. This could enable an open mashup of information that is at the heart of research. It also gives libraries a way to let in tools like the TAPoR Tool bar. For that matter we might see creative tools coming from out students as they fiddle with the technology in ways we can’t imagine.
The key difference between OpenSocial and the Facebook Platform is that the latter is limited to social applications for Facebook, as brilliant as it is. OpenSocial can be used by any host container or social app builder. Some of the other host sites that have committed to using is are Ning and Slide. Speaking of Ning, Marc Andreessen has the best explanations of the significance of both the Facebook Platform phenomenon and OpenSocial potential in his blog, blog.pmarca.com (gander the other stuff on Ning and OpenSocial too).
