Archive Zaro

Ralph Padilla, a DH (graduate) student that I work with, gave a talk today about a neat project of his own, Archive Zaro. The project is archiving video games from the Philippines. It has just launched.

This archive is part of an ongoing academic project on game preservation and digital heritage. It explores how stories, creative processes, and design histories of Filipino video games can be documented and studied without relying on commercial hosting or proprietary platforms. Each entry serves as both a cultural record and a research artifact, contributing to the broader understanding of how local game development reflects art, labor, and identity.

Declaration of Independence – First E-Text

Project Gutenberg and the Declaration of Independence

I came across a blog post about how Michael S. Hart, the founder of Project Gutenberg started in 1971 by typing the Declaration of Independence into the ARPANET and sending it to others. See 50 Years at Project Gutenberg.

Forty-Five Years of Digitizing Ebooks: Project Gutenberg’s Practices by Gregory B. Newby is a longer thing on the history of Project Gutenberg’s processes.

Hart passed in 2011. Gregory B. Newby just passed away this October. The Project, however seems to be in good hands with a foundation and board.

The Lives of Literary Characters

The goal of this project is to generate knowledge about the behaviour of literary characters at large scale and make this data openly available to the public. Characters are the scaffolding of great storytelling. This Zooniverse project will allow us to crowdsource data to train AI models to better understand who characters are and what they do within diverse narrative worlds to answer one very big question: why do human beings tell stories?

Today we are going live on Zooinverse with our Citizen Science (crowdsourcing) project, The Lives of Literary Characters. The goal of the project is offer micro-tasks that allow volunteers to annotate literary passages that help annotate training data. It will be interesting to see if we get a decent number of volunteers.

Before setting this up we did some serious reading around the ethics of crowdsourcing as we didn’t want to just exploit readers.

 

A Bored Chinese Housewife Spent Years Falsifying Russian History on Wikipedia

She “single-handedly invented a new way to undermine Wikipedia,” says a Wikipedian.

From Vice a rather funny story about how A Bored Chinese Housewife Spent Years Falsifying Russian History on WikipediaUser Zhemao wrote hundreds of linked articles in the Chinese version of the Wikipedia about fictional events, peoples and places in Russian history. Only recently did someone notice. It shows a vulnerability of such crowdsourced resources; a fabulist can create a network of consistent fictions that supporting each other look true.

GameStop, AMC and the Stock Market’s Wild Ride This Week

GameStop Stock Price from Monday to Friday

Here’s what happened when investors using apps like Robinhood began wagering on a pool of unremarkable stocks.

We’ve all been following the story about GameStop, AMC and the Stock Market’s Wild Ride This Week. The story has a nice David and Goliath side where amateur traders stick it to the big Wall Street bullies, but it is also about the random power of internet-enabled crowds.

Continue reading GameStop, AMC and the Stock Market’s Wild Ride This Week

MIT apologizes, permanently pulls offline huge dataset that taught AI systems to use racist, misogynistic slurs

Vinay Prabhu, chief scientist at UnifyID, a privacy startup in Silicon Valley, and Abeba Birhane, a PhD candidate at University College Dublin in Ireland, pored over the MIT database and discovered thousands of images labelled with racist slurs for Black and Asian people, and derogatory terms used to describe women. They revealed their findings in a paper undergoing peer review for the 2021 Workshop on Applications of Computer Vision conference.

Another one of those “what were they thinking when they created the dataset stories” from The Register tells about how MIT apologizes, permanently pulls offline huge dataset that taught AI systems to use racist, misogynistic slurs. The MIT Tiny Images dataset was created automatically using scripts that used the WordNet database of terms which itself held derogatory terms. Nobody thought to check either the terms taken from WordNet or the resulting images scoured from the net. As a result there are not only lots of images for which permission was not secured, but also racists, sexist, and otherwise derogatory labels on the images which in turn means that if you train an AI on these it will generate racist/sexist results.

The article also mentions a general problem with academic datasets. Companies like Facebook can afford to hire actors to pose for images and can thus secure permissions to use the images for training. Academic datasets (and some commercial ones like the Clearview AI  database) tend to be scraped and therefore will not have the explicit permission of the copyright holders or people shown. In effect, academics are resorting to mass surveillance to generate training sets. One wonders if we could crowdsource a training set by and for people?

Digital Synergies Launch Event


Today I gave a short talk at the Digital Synergies Launch Event. The launch included neat talks by colleagues including:

I showed and talked about Lexigraphi.ca – The Dictionary of Worlds in the Wild. This is a social site where people can upload pictures of text outside of books and documents and tag the words – text like tatoos, graffiti, store signs and other forms of public textuality.

Canadian Social Knowledge Institute

I just got an email announcing the soft launch of the Canadian Social Knowledge Institute (C-SKI). This institute grew out of the Electronic Textual Culture Lab and the INKE project. Part of C-SKI is a Open Scholarship Policy Observatory which has a number of partners through INKE.

The Canadian Social Knowledge Institute (C-SKI) actively engages issues related to networked open social scholarship: creating and disseminating research and research technologies in ways that are accessible and significant to a broad audience that includes specialists and active non-specialists. Representing, coordinating, and supporting the work of the Implementing New Knowledge Environments (INKE) Partnership, C-SKI activities include awareness raising, knowledge mobilization, training, public engagement, scholarly communication, and pertinent research and development on local, national, and international levels. Originated in 2015, C-SKI is located in the Electronic Textual Cultures Lab in the Digital Scholarship Centre at UVic.