Tuesday, April 22, 2014

Digital Preservation on Social Media: Some Issues

The preservation of photographs, metadata, comments, and other items posted to social media and social networks is a complex issue, and the archives world does not yet have standardized practices for handling these kinds of materials. Concerns facing the archivist include privacy, digital formats, site policies, and gaining access to the materials.

[A note on terminology: "social network" here refers to the kind of site where the goal, or a major component of the goal, is to collect all your friends like Pokemon. "Social media" is more amorphous, but refers to any site that has some kind of "friending" or "following" mechanism, and whose content is entirely or almost entirely generated by the users, not by the site staff. Facebook and LinkedIn are social networks. Twitter, Tumblr, and Instagram are social media services. Things like Flickr aren't quite either, but the questions of whether to preserve comments made to photos and how to carry out that preservation are related to social media/networking preservation.]

Archivists must consider the privacy rights not only of the creator or original poster, but also of the commenters. In a restricted-access model such as Facebook, users make posts and comments based on their assumptions (or their actual knowledge, in some cases) of who will have access to view their content. Making these materials available more widely than their original context may engender the wrath of privacy advocates and certain users; however, similar conflicts have occurred regarding the original poster's control over privacy settings - particularly in a system with granular (item-level) privacy settings, users have the ability, and sometimes the inclination, to make previously-restricted materials public. The general trend in behavior of social media operators/administrators/Terms of Service enforcement teams has been toward affirming the rights of the original poster over the 'entire work' comprising the original post plus comments. Anonymous commenters retain no rights; logged-in commenters generally retain the right to delete their comments, sometimes also to edit them, but they have no privacy controls. The conflict over privacy, particularly commenter privacy, comes from users' expectations, not legal concerns.

The Library of Congress's Twitter archive handles privacy by only collecting public tweets (Twitter users can choose to restrict their entire account, meaning that they must approve each new follower before that person can see their tweets.) Tweets are also embargoed for six months. But Twitter is a simple privacy use case - an account is either public or protected, and replies to tweets are separate tweets linked together by Twitter's software, not traditional comments appended to and dependent on an original post.

Occasionally, social media services make content available through standardized APIs, for which a repository (or a coalition of repositories, or some other entity) could develop a standardized tool that accesses the site through the API and saves files locally or to a server controlled by the repository, enabling archivists to automate digital preservation of material from those sites. Unfortunately, it's far more common for sites to use homebrew or proprietary APIs (meaning that a new tool would need to be developed for each site or obscure API) and to make only portions of content available through the API. Additionally, APIs are often rate-limited; someone accessing the site through the API can only make so many requests in a period of time before the service stops responding, until the clock resets. Web crawlers (like the one the Internet Archive offers as Archive-It) can help to fill the gaps left by APIs, but these are optimized for static web pages and don't have the ability to bypass login screens.

And these automated tools don't touch the issue of digital formats. Web image formats such as JPEG and PNG seem fairly long-lived, but the repository will still need a system that backs up the preserved files and checks them for corruption - and some method for transitioning copies of the original files to new formats if the original formats become obsolete. As for preserving comments and metadata, in the absence of a tool that utilizes an API or a web crawler, there would seem to be three options, each with drawbacks. Copying the information to a database controlled by the repository would be prohibitively labor-intensive. Saving each page locally as a PDF would create copies in a very stable format, at the expense of ease of searching and guaranteed loss of any dynamic features of the website. Saving each page locally as an HTML file may be the best option; if enough pages are saved in this way, at least some of their links may continue to work, and there is a chance that the dynamic features will be preserved.

A small trend toward appointing "digital executors" may aid archivists in preserving social media materials. [A digital executor is someone who knows how to obtain your passwords after your death; they're supposed to carry out your online last wishes.] Particularly if the digital executor also assumes control of other materials that might be donated to a repository, it may be possible to negotiate a donation of digital "papers" as well. Issues of format would still enter the equation, but archivists would understand what level of access they achieved, rather than having to make guesses.

In many cases, preservation from social media sources will be up to the users themselves; materials from these sites will only arrive at repositories when users have chosen to gather this information and include it with papers they donate to an institution. Archivists are - at least at present - unlikely to seek out the online presence of an individual and preserve it. Yet very few users carry out personal preservation work; their assumption tends to be that uploading something to "the cloud" is itself preservation.

Complicating the idea of archivists conducting the preservation work are social media services' own policies (many will automatically delete accounts after a certain period of inactivity) and the relatively ephemeral nature of the internet. Certainly, wave after wave of social media sites have failed to capture a sufficient userbase, shutting down and destroying whatever was posted to them. But consider what will happen when a Facebook or similarly popular and long-running service is supplanted by something completely new, and shuts down. When Geocities (in a way, a precursor to social media's focus on user-generated content) closed in 2009, it wiped all its content from the web, deleting huge amounts of user-generated content, though much was preserved through the work of multiple online groups. But if a social media site with privacy controls - particularly one that implements a non-public privacy setting by default - were to go dark, many of the techniques used to preserve GeoCities would fail.

That said, efforts like the GeoCities Project offer a suggestion for how to handle massive social media preservation efforts on the scale of "entire services." After acquiring as much of the content as possible before GeoCities' shutdown, Archive Team released a torrent. BitTorrent is more popularly associated with piracy of music, movies, and software, but its behaviors (including peer-to-peer transfers and the ability to download specific portions of the material the torrent points to) also make it ideal for granting access to massive quantities of minimally-processed archived digital material.

Archive Team. "GeoCities." Last modified April 1, 2014. http://archiveteam.org/index.php?title=GeoCities.

Greene, Kelly. "Passing Down Digital Assets." Wall Street Journal, August 31, 2012. Accessed April 21, 2014. http://online.wsj.com/news/articles/SB10000872396390443713704577601524091363102.

Internet Archive. "Archive-It - Web Archiving Services for Libraries and Archives." Accessed April 21, 2014. https://www.archive-it.org/.

Library of Congress. "Personal Archiving." Accessed April 20, 2014. http://www.digitalpreservation.gov/personalarchiving/index.html.

McNealy, Jasmine. "The Privacy Implications of Digital Preservation: Social Media Archives and the Social Networks Theory of Privacy." Elon University Law Review 3 (2012): 133-160.


Spiliotopoulos, Dimitris, Efstratios Tzoannos, Pepi Stavropoulou, Georgios Kouroupetroglou, and Alexandros Pino. "Designing User Interfaces for Social Media Driven Digital Preservation and Information Retrieval." In Computers Helping People with Special Needs, 581-584. Berlin: Springer, 2012.

No comments:

Post a Comment