Web Crawler: Nutch

Nutch is “open source web-search software. It builds on Lucene Java, adding web-specifics, such as a crawler, a link-graph database, parsers for HTML and other document formats, etc.” There is a Nutch Wiki with links to news, presentations and articles on it.

Nutch is basically a open Google-like engine that indexes an intranet (or the web) and gives you search capability. This sort of tool could be useful if there were ways to adapt it to discipline specific crawling.

Latent Semantic Analysis

LSA @ CU Boulder is a site at the University of Colorado at Boulder on Latent Semantic Analysis for education. The neat thing is they provide a web interface to different LSA tools. Could these techniques be used in research text analysis? Could we create them as web services?

A site they point to with a list of links to readings, projects and people is Readings in Latent Semantic Analysis, maintained by Lemaire and Dessus.

They also link to a Wired News article on LSA in education that explains how LSA can be used for automatic marking of essays, see Teachers of Tomorrow?.
Continue reading Latent Semantic Analysis

ATLAS.ti

ATLAS.ti! is a “Knowledge Workbench” for the qualitative analysis of texts, images, audio and video. It looks like a PC program that lets you annotate large quantities of materials for interpretation, coding, and clustering.

I saw this years ago, but it has matured and now handles multimedia. I should add that it is for sale, not free, though they have a trial version.
Continue reading ATLAS.ti

What is an electronic text?

I came across a thoughtful blog entry responding to somethng I wrote with Ian Lancashire about electronic texts and text analysis in WRT: Writer Response Theory ª Forms of Electronic Texts.

The author, Christy Dena, points out the focus on material characteristics (that an e-text is an electronic version of a written work etc.) and inconsistencies. To be honest I wasn’t trying to come up with a typology with a “continuity of variable.” I was trying to describe the variety of things we call e-texts. Time for a better definition and asking whether we want to use “electronic text” for anything that can be read and has/had an electronic form.

Unstructured Information Management Architecture (UIMA) from IBM

According to this Reuters article, Search concepts, not keywords, IBM tells business, (August 8th, 2005) IBM is releasing their UIMA SDK (Unstructured Information Management Architecture Software Development Kit) to developers as open-source.
According to an IBM Overview the UIMA provides tools for improving text searching through “analysis technologies, including statistical and rule-based Natural Language Processing (NLP), Information Retrieval (IR), machine learning, and ontologies.” Unstructured information includes not only text, but audio, video and images. This is thanks to Mike Rowse.

http://www.alphaworks.ibm.com/tech/uima/
Continue reading Unstructured Information Management Architecture (UIMA) from IBM

Buzz Engine: Online Analysis

According to a Globe and Mail story, Buzz cuts through the on-line rumour mill, (Jerry Langton, August 4, 2005, Globetechnology Section) Accenture researchers have developed a technology called the Buzz Engine that tracks topics through lists and blogs. It looks like it does something like the culture tracker we developed, graphing the relative frequency of keywords – real-time text analysis.

Here is a quote from Gary Boone, PhD: Weblog:

At Accenture Technology Labs, we have developed the next generation of search engine. Itís a kind of summary engine that focuses on online buzz or discussion. Online Audience Analysis is a buzz engine that interactively shows how much buzz there is on a given topic. You can search for topics of interest and see how much public attention that topic receives. Is anyone talking about the new Xbox? Are more participants in technology discussion sites talking about iPods or about Creative Zen Micros? Online Audience Analysis can show you.

Continue reading Buzz Engine: Online Analysis

Echelon doesn’t seem to work

A story about how the British intelligence services have been closing down al-Qaeda related web sites, Finger points to British intelligence as al-Qaeda websites are wiped out (from the The Sunday Times, July 31st, 2005), comments that automated electronic intelligence gathering systems like Echelon don’t seem to work. In other words text-analysis systems don’t work if people want to subvert them by using simple codes or spamming the net. See my previous posts on Carnivore Documents.
Does this mean it is unlikely to be helpful to students of textuality?