Sonic Word Cloud Project Part 4: Conclusions

This is the fourth and final blog post of a series in which I discuss and demonstrate my final project for the course I am taking on Digital History. In my first post, I discussed the origins and the purpose of my project, which consists of a method to sonify the most frequently used words in a text in order to hear their linear relations to one another over time. In the second post, I covered how I put together a python script that can automatically output MIDI and text-to-speech mp3 files for the 25 most frequently used words of a given text. In the third post, I showed how I combined these files together in music sequencing software and attempted to refine the end result. I also began to discuss some of the challenges and limitations of sonic word clouds as a form of textual representation. In this post, I will conclude with some thoughts weighing the overall usefulness and applicability of this method of sonification.

I first tested out this method of sonification on the 1916-1918 Robert Lindsay Mackay Diaries [1], and above is the final result. From the outset, my goal was to figure out whether this method to sonify historical sources could be at all useful for the purposes of 1) textual analysis, and/or 2) public textual representation. As I have discussed in my previous post, I have more or less ruled out the possibility of the latter (at least in the project’s current state) due primarily to the ‘glitch effect’; however, I argue that there are some contexts in which it could still potentially be useful for the former.

One way I attempted to show this was through an example of ‘targeted analysis,’ illustrating how one can work within the sequencing software to solo specific words and hear how their occurrences relate to each other in the text over time. With Robert Lindsay Mackay’s diaries, for instance, I explored the relation between event words such as ‘killed’ and ‘wounded’ and Othering terms for the Germans such as ‘hun’ and ‘boche,’ noting a spike in the latter words following a spate of ‘killed.’ This kind of analysis does not provide any definitive conclusions on its own, but could help to spark and/or guide new lines of inquiry in the analysis of a text. Conceivably, I could also make another version of the Python script that would allow the user to search for custom words in a text and output midi and mp3 files for it. This would allow for much more flexible analyses of the relations between the usages of different (and possibly more obscure) words in a text.

Screen Shot 2017-04-13 at 5.25.16 PM
The template project file for Ableton

I also attempted to develop this method in a way that made it relatively easy to apply it to different texts in a short period of time. In terms of the Python script, this was a success, as in its current form it quickly and painlessly produces midi and mp3 files from any text that’s fed into it. The stage where these files are combined in sequencing software using samplers, however, remains a bit more time-intensive. I have eased it a bit by creating a template project file for Ableton, containing the needed 25 Sampler tracks (each with a separate panning position). With this, the most time-intensive part is to load the audio files into the sampler tracks and make sure that the samples trigger at the right starting point. I would estimate that it now takes me about 15 minutes to carry out this method from start to finish on a given text, which is not too bad; compared to the almost immediate results of most visualization tools, however, this still takes more time than I would like.

I have tested out the sonic word cloud method on a variety of different historical texts, and in the end I’ve found it to be most useful for personal narratives that span a significant period of time. These include works written over time in close proximity to events that they describe (such as the Robert Lindsay Mackay journals) as well as narratives written after the fact (such as autobiographies). Above is a sonic word cloud created from  the 1789 autobiography of Olaudah Equiano a.k.a. Gustavas Vassa (c. 1745-1797). [2] This sonification is interesting for how it contains shifting patterns of frequent word usage that reflect the narrative of Equiano’s life. The first part of the sonification, corresponding to where Equiano relates his early life and his initial captivity into slavery, reveals a focus on difference, as well as the dichotomy of master and slave. The introduction of nautical terms (‘ship,’ ‘captain,’ ‘vessel’) hint at his journey in captivity across the Middle Passage, and their continued prevalence alongside ‘master’ and ‘slave’ reflect how he continued to labour on and around ships for a large part of his period of enslavement. The spike in ‘god’ at around the 4 minute mark reflects Equiano’s conversion to Christianity later in his life, while overall declines in the frequency of ‘master’ and ‘slave’ follow his emancipation. These are narrative elements that the sonification hints at in ways that could not be equaled by a visual word cloud, which represents the frequency of word usage in a more static and monolithic form.  [3]

Screen Shot 2017-04-14 at 3.27.48 PM
A visual word cloud of Olaudah Equiano’s autobiography in Voyant Tools

However, at the same time sonic word clouds also contain some of the same potential issues and dangers of visual word clouds. The biggest danger is seen in how it decontextualizes textual data. In a sonic word cloud, the most frequently appearing words are removed from their context and thus their semantic meanings become obscured. This can lead to potentially misleading representations of the contents of a text. For instance, as I discussed in my last post, “coy” was one of the most frequent words represented in the journals of Robert Lindsay Mackay, but after looking into it a bit I found out that this was in fact being used as an abbreviation for “Company.” I had to go to the original text to confirm this. By contrast, visualization tools such as Voyant have features such as “Contexts” that easily provide n-grams for any given search term, which help give a sense of how a given word is being used in the text. Thus one must be careful not to take the words represented in sonic word clouds at face value. Moreover, much as in the case of visual word clouds, if a word is used in different contexts with different meanings in a text, these meanings will be grouped together and represented singularly within a sonic word cloud.

Screen Shot 2017-04-14 at 4.01.52 PM
The term ‘coy’ viewed in the Contexts tab of Voyant Tools.

In spite of these limitations, the linearity of sonic word clouds still offer some useful advantages over visual word clouds as a tool of textual analysis, since they preserve certain aspects of the narrative of a text and can reveal relations between the usage of different words within it. However, as mentioned in my first post, there are other visualization tools, such as the “Trends” feature in Voyant, that can also plot the broad trends in the usages of different words in relation to each other on a graph. One of the only advantages of sonic word clouds over these visualizations is that they represent every single occurrence of a given word in a text, rather than an approximation of its relative frequency over time. This would theoretically allow one to hear the outliers in the usage of the most frequent words, and their relation to the usages of other words, in addition to the broader patterns. At the same time, though, the time it takes to assemble the sonic word cloud (~15 minutes) and then listen to it (~3-5 minutes) significantly limits its immediate usefulness relative to forms of visualization.

Screen Shot 2017-04-14 at 5.23.37 PM
Voyant Tools‘ Trends feature, applied to the Robert Lindsay Mackay journals.

Overall, much like other digital humanities tools, sonic word clouds cannot stand on their own. That is to say, their usefulness is found not in the questions they can answer regarding the linguistic patterns in a text, but rather in the questions they can potentially raise to guide further inquiry and analysis. As it stands, however, the time costs involved in using this method for textual analysis makes its overall usefulness in comparison to other tools somewhat doubtful.

The python script (midifreq.py), the obo.py library it relies on, the template project file for Ableton and the sample text with the corresponding MIDI and mp3 files can be downloaded here:

Sonic Word Cloud Project (Google Drive)

[1] Mackay, Robert Lindsay. The Diaries of Robert Lindsay Mackay. http://www.firstworldwar.com/diaries/rlm.htm

[2] Equiano, Olaudah. The Interesting Narrative of the Life of Olaudah Equiano. London: Olaudah Equiano, 1789. http://www.gutenberg.org/files/15399/15399-h/15399-h.htm.

[3] Harris, Jacob. “Word Clouds Considered Harmful.” NiemanLab (http://www.niemanlab.org/2011/10/word-clouds-considered-harmful/.

Sonic Word Cloud Project Part 3: Bringing it All Together in Ableton

This is the third post in a series where I explain and demonstrate my final project for the course I am taking on digital history. In my first post, I discussed the origins and the purpose of my project, which consists of a method to sonify the most frequently used words in a text in order to hear their linear relations to one another over time. In the second post, I covered the process of how I put together a Python script that can automatically output MIDI and mp3 files for the 25 most frequently used words of a given text. In this post, I will show how I combined these files together in music sequencing software (in this case Ableton Live), demonstrate what the end result sounds like, and also begin to discuss some of the challenges and limitations of this form of representation.

Once I had refined the Python script to the point where it could easily take any text and produce MIDI and text-to-speech mp3 files for its most frequently used words, all that remained was to bring all of these files together in music sequencing software. To do this, I used the software Ableton Live. (theoretically though, it’d be possible to do the same thing in other sequencers, including free ones such as LMMS). As I touched upon in my first post, my method of sonifying the midified data relies on samplers. A sampler is basically a digital instrument that uses MIDI notation to play any audio sample that is loaded into it. The sample can also be manipulated in various ways within the sampler. As mentioned previously, the text that I first tested this method on was the journals of Robert Lindsay Mackay. So, the first step was to load twenty-five sampler tracks (using Ableton’s built-in Sampler) into a new project file in Ableton. This can be done very quickly by loading one sampler and then copy & pasting it 24 times.

Screen Shot 2017-04-13 at 5.25.16 PM.png
Step 1: Twenty-five Samplers are loaded into Ableton

From there, I needed to load the mp3s of the top 25 words in the journals that were produced by the Python script into the samplers, and then rename each track to indicate the word it corresponded to. This was the most time-intensive part of putting all of it together.

Screen Shot 2017-03-10 at 5.14.50 PM
The mp3 for the word “arras” loaded into a Sampler.

As you can see illustrated above, the audio file for each word contains a short period of silence before the word is actually spoken. If left as is, this would mean that the MIDI data of each occurrence of the word within the text would not be accurately represented by the sonification, as these gaps make it so that that the spoken words will not correspond to the MIDI notes that trigger them. This can be solved by moving the marker on the left to cut the silence preceding the word out of the sample. This had to be done for all 25 of the Sampler tracks, which takes a bit of time (different words/sounds will have different optimal starting points, which have to be determined by ear).

Screen Shot 2017-03-10 at 5.15.00 PM
The “arras” sample, altered to play the audio of the spoken word as soon as it’s triggered by MIDI

Once all of the sampler tracks were loaded with their respective audio files, the next step was to bring in all of the corresponding MIDI files. Since I had loaded the audio files into the samplers in alphabetical order, it was fairly simple to just drag and drop all 25 of the MIDI files at once into their corresponding sampler tracks.

Screen Shot 2017-03-10 at 5.11.50 PM
A view of the Robert Lindsay Mackay journals’ MIDI files loaded into their corresponding Sampler tracks. You can also see previews of the MIDI notation contained in each MIDI file, represented by the small black blocks.

This was a good start, as all of the MIDI files overlaid each other without any major problems or discrepancies. However, there were still a few issues to work out. Most glaringly, the overall length of the combined MIDI files was about 85 minutes long if played at 500 BPM (the tempo that I had defined in the Python script). This was obviously way too long to be useful, and even if I played it at Ableton’s maximum of 999 BPM, it wouldn’t be sufficient to bring the overall time down to the ~3 minute time frame that I thought to be a reasonable place to start. Luckily though, Ableton has a feature allowing you to condense the length of a MIDI file by half with a single button. However, each of the MIDI files produced from the journals has a slightly different length, depending on how closely the word it represents appears to the end of the text. This meant that in order to condense all of the MIDI files at once consistently, they had to all become the same length. Fortunately, MIDI clips can also be consolidated together to become the same length in Ableton. So once all of the MIDI files were consolidated, it was fairly simple to condense them all at once down to the point where they could be played in the span of about two and a half minutes.

Screen Shot 2017-03-10 at 5.15.31 PM
The journals’ MIDI notation, consolidated and condensed down to about 2.5 minutes.

With all of the twenty-five most frequent words of the journals represented through MIDI files mapped to text-to-speech samplers, and reduced down to a more digestible length of time, I could now finally listen to what the resultant sonic word cloud sounded like. The hope was that some broad patterns of Robert Lindsay Mackay’s thoughts and experiences over the two years of his service overseas could be revealed through this linear sonic representation. Pressing play, this is what I heard:

My first impression was that it sounded much more chaotic than I had hoped. It was relatively difficult  to distinguish between all of the different spoken word samples being triggered throughout, even when knowing beforehand what words to listen for. Moreover, the frequency with which the samples are triggered by the MIDI notation produces series of overlapping triggers, wherein just the beginning of the sample is re-triggered 10-20 times in the span of a couple seconds before allowing the rest of the sample to be played. This results in a glitch-type effect in many parts of the audio file, heard most dramatically with a rapid-fire series of the word ‘killed’ at 1:20 (which corresponds to a part  halfway through the journals where Mackay recounts the fates of all the men in his battalion).

There were multiple ways by which I worked to lessen some of these issues. First, it was clear that if all of the words were to be included in the sonification, the BPM would have to be reduced in order to lengthen the final audio file and thus give the MIDI notations more room to breathe. A balance must be struck here between intelligibility and usefulness. I decided that 4-5 minutes was the maximum length that was reasonable  to listen to for this kind of representation. The necessity to listen to a sonified representation over this period of time unfortunately makes it less immediate than visual forms of representation, limiting its relative usefulness as a means of textual analysis.

Screen Shot 2017-03-10 at 5.34.30 PM

Second, I introduced stereo separation and speech enhancement EQ (equalization) in order to increase the comprehensibility of the word cloud and make it more listenable. Stereo separation involved panning each sampler track from left to right so that each track/word would have its own separate place in the ‘soundstage’ (as roughly illustrated above). In addition to making the soundstage less cluttered in the centre, this would help the listener to know ‘where’ to listen for a given word as it’s triggered over time. I started by panning the tracks alphabetically, since that was the simplest option. I then created a template project file in Ableton with 25 empty sampler tracks, each with a different pan position, in order to save time when working with future texts. However, this panning could also be altered to order the words into different groupings and thus aid with analysis; for instance, ‘event words’ such as killed, wounded, and work could be panned to one side, with ‘place words’ such as Arras, road, and place panned to the other side, in order to help gain some insight into the relations between these word groupings.

Another way to de-clutter the word cloud was to mute words that did not add to analysis of the text or were otherwise unnecessary. Some of these choices were pretty clear-cut; for instance, “officers” did not necessarily need to be included alongside “officer” (ideally, there would be a way to combine the two together), and words such as “just” do not seem to provide much insight into the content of the journals. Other cases are less straight-forward, though, and involve some subjective privileging of certain kinds of information over others. For instance, in any journal instances of “day” and “night” can be expected to make up a large part of the text; can they be safely left out of the final word cloud? Or would doing so risk literally silencing a part of Mackay’s experiences and what he considered important to record? This would be more of an issue if this method were to be used for representation (e.g. in a public history display) than for textual analysis , however, and so we will set it aside for now. Another issue came with the word ‘Coy,’ which is actually used as an abbreviation for “Company” in the text (which I first realized when I listened to the audio tracks and noticed that ‘Coy’ was almost always triggered at nearly the same time as ‘Arras’). I ended up deciding to mute the ‘officer,’ ‘day,’ ‘night,’ ‘morning,’ ‘coy’ and ‘just’ tracks for the time being.

Next, I changed the voice option in gTTS from the default (American) accent to the British accent. Beyond being slightly more fitting for sonifying the journals of a British soldier, I also found this voice to be a bit easier to understand.

With these alterations combined, there were noticeable improvements in regards to the comprehensibility of the audio, as can be heard in the above audio file. The stereo separation in particular helps to make it easier to understand, in conjunction with the slower tempo which gives the MIDI notation a bit more room to breathe. This makes it easier to get a sense of the broad patterns of word usage throughout the text in the sonification, which could potentially raise questions that may help to guide closer examination and analysis of the original text.

One can also carry out more targeted analyses by soloing a smaller amount of words, in order to gain greater insight into the relations between their occurrences in the text. For instance, in the above audio clip we hear the relations between occurrences of event words such as ‘killed’ and ‘wounded,’ ingrouper terms such as ‘men,’ and derogatory terms for the Germans such as ‘boche’ and ‘hun,’ which can possibly provide insight into how Mackay’s experiences of war affected his Othering of the German enemy. For instance, in the clip we can hear a spike in the incidences of ‘boche’ and hun’ following the rapid-fire series of ‘killed’ halfway through (at around 1:00). Since there’s a lot less words to keep track of here, we can afford to play the MIDI clips at more than twice the BPM of the previous comprehensive word cloud, condensing the overall length of the sonification to only two minutes. This kind of targeted analysis to uncover relations between usages of different words throughout a text is one way that this kind of sonification could potentially be used as an analytical tool.

However, other issues remain unresolved. The main one is the glitch issue that comes from word samples being triggered multiple times in the span of a second. As far as I know, the only real way to solve this (at least for someone with my relatively limited coding skills) is just to reduce the master tempo to the point that the overall sonification is 40-50 minutes long. At that point, you might as well just read the original source. This issue prevents this kind of sonification method from becoming an adequate form of textual representation (e.g., in the field of public history) at present; however, it is my opinion that the glitch issue does not entirely negate this method’s (albeit limited) potential for textual analysis, as the glitch effects still give a clear indication that the word in question is occurring in the text with rapid frequency.  The problems do not end there, however; in my next blog post, I will discuss some of the broader potential issues with how this method represents textual data, before ending with some final thoughts on the overall upsides and downsides of this experiment.

Sonic Word Cloud Project Part 2: The Python Script

This is the second of a series of blog posts in which I explain and demonstrate my final project for the course I am taking on Digital History. In my first post, I discussed the origins and the purpose of my project, which aimed to develop a method to sonify word usage data in texts. In this post, I will begin to explain the process by which I started to put it together, and some of the challenges I faced.

As my planned method for sonifying word usage data in a text would rely on the MIDITime Python package, I knew that I would have to gain some familiarity with Python in order to use it for my purposes. In the end, the bulk of time I spent on my project went towards coding a Python script that would take a text, determine the 25 most frequently-used words in it, return the data of each time the word occurs in the text, and then convert that data into MIDI files with the MIDITime module. So I started out by going through a bunch of lessons on the site Codecademy to learn the core concepts of the Python language. Then I adapted a couple of the Programming Historian lessons on Python, as I have discussed in a previous blog post. This provided me with a script that could take a given text file and print a series of its word-frequency pairs. I tested it out on the Robert Lindsay Mackay diaries, and after appending some stopwords (including the letters of the alphabet and some numbers), it worked. Then, just to make sure, I tested it on another text, a 19th century piece of travel literature entitled Journal of an African Cruiser. This is the result I got:

Screen Shot 2017-03-05 at 1.11.47 AM
Strangely, ‘\xe2’ was listed as the most frequent word. After looking around on the internet a bit, I found out that xe2 is the unicode for the character ‘â’ (which I confirmed by searching for it in the original text). By looking on StackOverflow, I found out how I could replace all occurrences of that character with ‘a’ by adding on to the code. This was a less messy solution than the alternative, which would have been to just append ‘\xe2’ to the list of stopwords (and thus potentially break up any words in which the character appears).

Screen Shot 2017-04-12 at 7.49.59 PM.png

From there, I needed to elaborate on the script to only produce the top 25 word frequency pairs, instead of all of them. This was fairly simple to accomplish using Python’s slice function.

Screen Shot 2017-04-12 at 7.49.28 PM
The first part of the python script. Line 17 slices the list of word-frequency pairs to return only the twenty-five most frequent words of a text.

Then I needed to code a function that would enumerate each word’s chronological occurrences in the text as a separate list, which it would then convert into MIDI data with the MIDITime module and output as MIDI files. This was the hard part. Luckily, I happened to stumble across a list comprehension on another StackOverflow thread that could take care of the main part of what I needed:

[i+1 for i, w in enumerate(fullwordlist) if w == tw]

This list comprehension basically takes a given word, checks the text for each time it occurs within it, and then returns a list of these occurrences expressed as values out of the total word count. I next had to figure out how to code a function that would both perform the list comprehension on all of the 25 top words in the text, and then convert the list of values into a form that the MIDITime Python module could read and then convert to MIDI notation.

Screen Shot 2017-04-12 at 11.43.52 PM.png
This is the code of the MIDITime module.

Basically, the MIDITime module requires four values to convert data into MIDI notes. Every note must have its time, pitch, velocity (which determines the amplitude/volume of the note) , and duration defined within the module. The time value is the most important one here, as it is the occurrence values of each word that will determine when the notes are played in relation to each other. The rest of the values can remain static: the pitch can stay at 60 (Middle C), the velocity can stay fixed at 127 (the maximum), and the duration can be defined as 200 beats across the board (since the occurrence values can go up to the final word count, the MIDI files will be very long and must thus be played at a high BPM/tempo). The challenge then,was to turn the lists of occurrences produced by the list comprehension into lists of MIDI notes that the MIDITime module could process.

Screen Shot 2017-04-14 at 8.11.45 PM
The first version of the getmidilist() function

Without getting too bogged down in the details, the getmidilist() function in the above piece of code does just that, making use of for-loops to carry out the enumerating list comprehension on each word as well as a midify() function that places the resultant occurrence values into lists that follow the MIDITime format before printing them as strings. This was one of the hardest parts of the script to figure out how to code, and it took a lot of trial-and-error to work out the logic of the for-loops. I tested it out on the Robert Lindsay Mackay journals, and eventually, I got it to successfully print out a list of MIDITime notes for the occurrences of each of the top 25 words in the text.

So now I had the ability to produce MIDITime notes for the most frequently used words in a text. From there the question became, how do I plug these lists of notes into the MIDITime module and actually convert them into MIDI files? At first, I thought I would have to just copy & paste the lists manually into the MIDITime script. But it quickly became apparent that it would take way too much time to do this 25 times over to produce MIDI files for all the top 25 words in a single text. I soon realized that I would have to figure out how to work the MIDITime code into my script, so that it could automatically read the list of notes for each word and output a MIDI file for it named after the word. This added another layer of complexity to the script, and figuring out how to accomplish this took many more hours of trial-and-error and frustration.

Screen Shot 2017-04-15 at 5.54.59 PM
The final version of the midioutput() function, with the MIDITime code incorporated into it

In the end though, with some help from StackOverflow, I eventually got it to work. Among other things, I found out that I had to modify the midify() function to return the list of notes as a list rather than a string, so that the MIDITime code could read it directly. The final version of the midioutput() function produces lists of notes for the occurrences of each top word in a text, and then uses the MIDITime module to output midi files for all of the top 25 words in the text, placing these files into their own folder in the directory of the original text file.

Screen Shot 2017-04-13 at 12.45.25 AM
The MIDI files produced from the diaries of Robert Lindsay Mackay via the midioutput() function in the script

With this, the most fundamental part of my project was now complete. I could now easily take any text, run it through my script and automatically produce MIDI files for the occurrences of the 25 most frequently used words in it. Next, I needed to find a way to produce audio files of each word being spoken (so that I could load them into to samplers and map the MIDI data to these samplers later on) . I decided that using text-to-speech software would be the quickest and easiest way to do this. At first I thought I would have to use it manually by entering each word into text-to-speech software one at a time to create separate audio files, but then I became aware of a Python module called gTTS (Google Text to Speech). It uses Google’s Text-to-Speech software API to output mp3 files for any text that’s defined within the module in Python. With this module, I could simply add on a piece of code to my script that would automatically create audio files for each of the top 25 words in a given text, and output the mp3s to their own folder in the directory of the text. This would save me a lot of time.

Screen Shot 2017-04-13 at 1.05.01 AM
The final part of my Python script. Google Text-to-Speech software is used to output mp3 files of the 25 most frequently used words in a text.

The above piece of code uses the gTTS module to do just that. For each word in the list of the 25 most frequent words of the text, it runs it through the Google Text-to-Speech module, and names each mp3 file after its corresponding word. It checks whether a file with that name already exists (using os.path), and if not it creates the file using the gTTS module.

It works fairly well and definitely saves a lot of time, but there are some downsides to using this module. For one thing, the audio quality of the output files is very low, consisting of mono 32kbps mp3s. This means that the quality of the final product is hindered significantly from the outset. Moreover, while the gTTS module offers different options regarding languages and even accents, it only features a female voice with no option to select a male voice. For a sonification of the journals of Robert Louis Mackay, this is perhaps not the most fitting voice. But due to its convenience, I ended up deciding to test the sonification with this method anyway, to at least figure out how well it works before exploring other potential options.

The Python script was now essentially complete. The last thing I did was to make it so that, instead of the text file being defined within the script itself, the user inputs a text file to analyze upon running the script. This made it easier to test out the script on different texts without having to get into the code every time. The final script can be downloaded here along with obo.py library from the Programming Historian that it relies on. (For it to work, the MIDITime python package must also be installed).

Screen Shot 2017-04-13 at 1.34.39 AM
The beginning of the final version of the script.

Now that I had a working script that could easily produce the needed MIDI and mp3 files from any text file, all that remained was to combine all of these files together within music sequencing software and hear what it sounded like. That process and its challenges will be explained in my next blog post.

Sonic Word Clouds: an Experiment with Data Sonification (Introduction)

This is the first of a series of blog posts in which I will explain and demonstrate my final project for the course I am taking on Digital History. In this post I will discuss the origins and the purposes of my project. For my project, I decided to experiment with a method of sonification as a means to analyze and represent historical data. More specifically, the idea was to develop a way to represent the most frequent word usage in texts through audio. For lack of a better name, I have decided to call this project a Sonic Word Cloud.

Screen Shot 2017-04-12 at 12.17.57 AM
 

Sonification refers to any method that is used to represent data through sound. So why would historians potentially want to sonify their historical data? Historian Shawn Graham has commented that by transforming data about the past into an aural form of representation, we make it new and unfamiliar. There is value in reshaping textual data into a form that we can approach with a new set of eyes (or ears), in order to analyze the data and ultimately find new patterns or meanings within it. This is true of visual representations, as exemplified by digital tools for textual analysis such as Voyant Tools that have by now become relatively commonplace. But it is perhaps even more true of sonic representations, which are even more unfamiliar to us at this point; as Graham notes, with forms of sonification “we begin to see elements of the data that our familiarity with visual modes of expression have blinded us to.” [1]

Screen Shot 2017-04-12 at 12.05.05 AM
A word cloud representing the combined works of Shakespeare (via Voyant Tools)

Word clouds are one visual form of textual representation that have become very common in text analysis tools as well as throughout the Internet as a whole. They take a text or a corpus of texts, determine the most frequently used words within it, and then represent these words clustered together with the size of each word corresponding to its relative frequency in the text. As a basic form of ‘distant reading,’ word clouds can help historians to quickly get a sense of the broad patterns of word usage within a given text or corpus.

However, word clouds also have significant limitations. First, they do not provide any indication of where words appear within the text. If, for instance, a word is used only in one small part of the text, then a word cloud could potentially give a misleading impression that the word is used throughout the text. Second, they do not offer any sense of how the words it represents are used in relation to each other within the text. Third, they do not provide a means to understand semantic meanings of different word usages; if a word is used in different contexts with different meanings (e.g., “mine” as a process of extracting ore, vs. an explosive device, or simply a possessive adjective), these meanings and usages will be grouped together as one word in the visual representation. Most pressingly, though, word clouds have attracted criticism for how they decontextualize textual data that other methodologies (such as n-grams or mapping) could illuminate more effectively. In a word cloud, the linear narrative of a text is reduced to a static representation that, some argue, perhaps obscures more than it reveals. [2]

Essentially, the question that drove me to plan my project was: how different or how useful might it be to represent the word usage data of a text linearly through sound? The website The Programming Historian features a lesson in which Shawn Graham demonstrates how to (among other things) use a Python package called MIDITime to convert different kinds of time-series data into MIDI files that can be mapped to instrumentation in music sequencing software. As an example, he shows how he converted topic modeling data from the entries of John Adam’s diaries into MIDI notation, which he then mapped to different instruments in Garageband. As Graham emphasizes, the choices of how to represent data through sound in this fashion are revealing of how we organize, privilege, reduce and transform information as historians. [3] However, the end result of this representation is not immediately intelligible to the listener; one may decide to, say, represent a topic concerning war in a text with the sound of piercing trumpets, but a listener has no means to truly understand the representation on its own without a separate guide. Therefore, this particular kind of sonification seems more useful as a kind of reflective process than as a practical method of textual representation or analysis.

Screen Shot 2017-04-12 at 12.31.59 AM.png
A view of Shawn Graham’s sonification of topics fitted to John Adams’ diaries in Garageband (via The Programming Historian)

My project takes inspiration from the Programming Historian lesson, but adapts and elaborates on it in other ways. The idea was as follows: to develop a method to take a given text, determine the top twenty-five words that appear within it, and convert the linear data of each time each word occurs in the text into MIDI notation with the MIDITime python package. From there, I would bring the MIDI files into music sequencing software, where they would then be mapped to samplers.

Screen Shot 2017-04-12 at 12.39.57 AM
A digital sampler viewed in the sequencer Ableton Live

A sampler is a kind of digital instrument that can play and manipulate any audio file that is fed into it. Each sampler in the sequencer would then have a text-to-speech audio file of one of the top twenty five words loaded into it. The MIDI file for each word would then tell its corresponding sampler when to trigger the spoken word. With all of the samplers and MIDI data combined, we could then hear a sonic, linear representation of word usage for the most frequently used words in the text in a way that is (in theory) readily intelligible to the listener.

Theoretically, there could be significant advantages to representing textual data in this fashion as opposed to a visual word cloud. First, we would be able to hear linear, temporal patterns in word usage, getting a better sense of where these words appear within the narrative of the text. Second, with all of the data combined, we could be able to hear how different words are used in relation to each other in the text, and how these relations change over time.

Screen Shot 2017-04-11 at 11.37.44 PM.png
Voyant Tools’ “Trends” visualization of the Shakespeare corpus

Practically speaking, this would perhaps be nothing too new; websites such as Voyant Tools offer graphs that can visually represent broad trends in the relative frequency of different words throughout a text or corpus. However, one difference with this proposed sonification method would be that the listener could be able to hear not only the broad trends, but also the outliers in word usage. This could potentially aid in guiding closer reading of a text.

Screen Shot 2017-04-12 at 12.48.54 AM.png
Robert Lindsay Mackay

As the potential benefits of this kind of sonic representation are mainly found in its linearity, I thought that it could be particularly useful for representing and/or analyzing texts that have been written over a significant period of time, such as journals. In light of the present centennary of the First World War, I thought it fitting to select the 1916-1918 diaries of a British soldier named Robert Lindsay Mackay [4] to be the text that I would first test this method on. From there, I just needed to figure out how to make this concept a reality, in order to determine its viability as a means of textual analysis and/or public textual representation (i.e. in a public history context). In my next blog post, I will chronicle the beginning of this process and its challenges.

The other installments of this series can be found here:

Sonic Word Cloud Project Part 2: The Python Script

Sonic Word Cloud Project Part 3: Bringing it all Together in Ableton

Sonic Word Cloud Project Part 4: Conclusions

[1] Graham, Shawn. “The Sound of Data (a gentle introduction to sonification for historians).” http://programminghistorian.org/lessons/sonification.
[2] Harris, Jacob. “Word Clouds Considered Harmful.” NiemanLab (http://www.niemanlab.org/2011/10/word-clouds-considered-harmful/.
[3] Graham, “The Sound of Data (a gentle introduction to sonification for historians).”

[4] Mackay, Robert Lindsay. The Diaries of Robert Lindsay Mackay. http://www.firstworldwar.com/diaries/rlm.htm

Exploring with Paper Machines

For my first in-depth look at a research-oriented digital history tool, I decided to try out Paper Machines, an add-on for the research/citation tool Zotero. The creators of Paper Machines promote it as a textual analysis tool that allows scholars to conduct large-scale ‘distant reading’ without requiring much technical knowledge. This sounded good to me.

screen-shot-2017-02-09-at-11-24-53-pm

Paper Machines

Much like some other prominent digital humanities tools, Paper Machines allows you to feed it a sizeable corpus of texts, which it then extracts and analyzes. From there, you can choose between a set of visualizations, such as word clouds, N-grams, phrase nets, topic modeling, etc. which will potentially allow you to draw broad conclusions about the content of these texts without having to actually read them all. In the case of Paper Machines, you collect these different texts and combine them into a corpus using the Zotero interface. I haven’t used Zotero much myself, but I’d imagine that this feature would make Paper Machines a big draw for those that already use it regularly.
After getting both Zotero and Paper Machines installed, the first step is to collect a corpus of texts for it to analyze. From doing a search on archive.org, I compiled 135 English texts concerning eugenics, dating from the early 1880s to the mid-1930s. Using Zotero, it was relatively quick and easy to collect all of these texts, as it saves all of the available metadata (including the date of publication) automatically, as well as a full-text pdf. Ultimately, my goal was to find out what kinds of insights Paper Machines could provide into the broad historical features and trends of eugenics discourse.

corpus
A view of the collection in Zotero

One thing I didn’t anticipate (though I probably should have) is how long it takes for the software to process the corpus. It took more than an hour for Paper Machines to extract the texts before they could be visualized. Then I selected the word cloud option, to get a basic sense of the most frequently used words in the corpus. A new tab opened up on my browser, a blue progress bar slowly filled up, and then… nothing. Just blank white space. I selected another type of visualization and got the same result. Eventually, I discovered that if I export the visualizations to my computer, and then load the html file from there, then I can get them to show up. I was able to access the word cloud through this workaround. But I still haven’t been able to get any visualizations to display in Zotero itself. ( apparently, this isn’t an isolated issue.)

eugenics-word-cloud
A word cloud visualization of the eugenics corpus

Just by looking at this visualization, one can start to get a sense of the key ideas that were employed in the discourse of eugenics: concepts such as heredity, population, race, class, gender, normality, children, motherhood, and conceptions of variously diseased or ‘defective’ individuals (with a particular focus on the ‘feeble minded’ or ‘insane’). We can also possibly get a sense of the ways in which eugenicists sought to legitimize their theories: common keywords such as ‘science,’ ‘fact,’ ‘nature,’ ’medical,’ ‘knowledge,’ ‘moral,’ ‘development,’ and ‘progress’ suggest how the enactment of eugenic principles or policies was framed as being natural, highly scientific, and morally imperative to human progress.
However, it is important to note that we cannot draw any firm conclusions from this word cloud alone. For one thing, it gives no indication of the contexts in which each word is used, and thus the full semantic meaning of some words is impossible to parse. For example, it is unclear whether the keyword “law” might be used to describe the abstract principles of eugenics that various authors argue for in the texts, or instead the actual eugenic policies (namely sterilization laws) that were enacted in different countries in the early 20th century. Second, the cloud does not indicate how concepts are related to each other in the texts. One can speculate (as I have in the previous paragraph), but there is no way to confirm or contradict these speculations using the word cloud alone. Therefore, the true value of this visualization tool is found not in the questions it can answer, but instead in the questions that it can raise for the researcher. It can potentially be very helpful in formulating research hypotheses, facilitating initial keyword searches and guiding the close reading of the texts.

screen-shot-2017-02-09-at-11-56-03-pm
The phrase net for “[x] of the [y].” Bolder lines indicate phrases that appear more frequently in the corpus

Another useful type of visualization that Paper Machines offers is the Phrase Net. After choosing a connecting word or phrase, such as “the” or “is,” the resulting Phrase Net will visually represent the most common phrases in the texts that follow the defined formula. For example, we can look at the phrase net for “[x] of the [y]” and get a sense of some of the anxieties of the eugenicists: such as the “population of the world,” “survival of the fittest,” “fertility of the unfit,” “size of the family,” “inheritance of the insane,” and “welfare/improvement/future of the race.” Their chilling solutions to these supposed problems are also highlighted by the phrase net: “control of the feeble,” “elimination of the unfit,” and most plainly, “removal of the testes/ovaries.” The phrase “duty of the state” suggests authors’ political arguments for the enactment of these eugenic sterilization policies.

Much like the word cloud, however, the phrase net is more useful for sparking or guiding lines of inquiry than it is for providing definitive answers. Besides the lack of wider context for the phrases, one of its significant weaknesses is that it does not include the temporal dimension of the data. This means that if, hypothetically, a particular concept or phrase had only become common currency in eugenics discourse around 1920, a phrase net could give the misleading impression that it was a feature of eugenics debates from the beginning.

topic-model-model
A topic model of the corpus

This is where Paper Machines’ Topic Modeling visualization adds another layer or two of sophistication. By determining which topics occur most frequently in the corpus, and plotting their occurrences on a time graph, one can get a better sense of the broad historical trends in a given body of texts. In the above topic model of the eugenics corpus, for instance, it can be noted that the explicitly racial topics (‘race, content, journal,’ ‘families, breed, inferior,’ and ‘decline, race, vice,’)  seem to predominate in the period from around 1900-1905, before being increasingly complemented by topics appearing to focus on nationalism, Malthusian population anxieties, and the mental fitness of individuals (‘science, mental, nation,’ ‘birth, rate, death,’ and ‘state, world, nation’). This may provide a sense of some of the shifting developments in eugenics discourse. Moreover, unlike in the other visualizations that Paper Machines offers, here you can actually see which specific texts include a topic in a given period by clicking on a point in the model, making the visualization’s representation of the data much more transparent to the user. The topic model also highlights significant gaps in the corpus, revealing which time periods are underrepresented.

Overall, despite some significant technical issues , Paper Machines is a relatively easy-to-use piece of software that offers historians helpful tools to aid in their research. Its integration into the Zotero interface makes it fairly easy to compile a large body of digitised texts in a short amount of time. Its visualizations can offer insights into broad discursive trends, and can thus help spark new lines of inquiry. The results of these various visualization tools should always be interpreted carefully (there is always the danger of positivism), but the opportunities they present for aiding in the analysis of a large body of texts are too useful to pass up.

Review – Musical Passage

A visitor to the site is first greeted by the calming sound of waves crashing steadily against a shoreline. Out of this soundscape, there slowly emerges a rhythmic melody played on an unfamiliar instrument (a mbira dzavadzimu, as I’m soon to learn). A flowing, repeated melody, its syncopated rhythm and recurrent tones are entrancing. The page shows a faded map of Jamaica, with the caption “Musical Passage: a Voyage to 1688 Jamaica.” Along the sides of the page can be seen parts of what looks like early sheet music. This all sparks some compelling questions: who originally composed/played this music? Is it really from 1688? How did it manage to survive in written form? And, most intriguingly of all, could this really be what it might have sounded like?

intro_1
Musical Passage

These are the questions that the digital history project Musical Passage aims to both invite and answer as best it can. Its approach is open-ended, engaging, and intuitive. The website is devoted to the preservation and interpretation of a single unique source: an excerpt from Hans Sloane’s 1707 Voyage to the Islands, which happens to contain the earliest transcription of African music in the Caribbean. Beyond the actual source itself, the site also offers audio recordings of different performances of these transcriptions. It features three songs, with titles referring to different African ethnic groups, that were likely played by enslaved musicians brought across the Atlantic via the Middle Passage. It is therefore an invaluable source for attempting to get a sense of the complex cultural blending that resulted from the African diaspora in the slave trade. It also provides valuable insight into the ways in which enslaved African people steadfastly preserved their cultural heritage, in spite of the displacement and oppression they had been subjected to.

screen-shot-2017-01-26-at-7-04-21-pm

Instead of giving a lengthy background on the context of the source and the music beforehand, the site instead opts to take the user directly to the source itself after a short introductory text. In doing so, the user is given the opportunity to see for themselves what the source says about these pieces of music, before exploring the pieces and pursuing the questions that the text raises for them in a non-linear fashion. Certain words and phrases are highlighted in red on the pages; if the user hovers over them, specific contextual or interpretive questions appear on the side (e.g. “Who played this music?” or “What was Jamaica like in 1688?”). If the user is curious about any of these questions, they can click on these phrases to bring up a little write-up that provides more historical context to the source. If, instead, they want to get the full story behind the source all in one place, they can simply go to the menu sidebar and click on “Read.” But to me, it felt much more engaging to be dropped directly into the source from the start, and have questions and issues organically come up from the text. It allows the user to explore the value, context and limitations of the source in a more interactive way, and so I think it’s a very effective way to present the background information.

screen-shot-2017-01-26-at-7-05-11-pm

The site’s greatest value, though, is in its recorded performances of the pieces included in the source. Music is an essential part of culture; unfortunately for historians of this period, though, it is also fluid and highly interpretive, living first and foremost in performance. With these transcriptions only conveying approximations of rhythm, melody and some (very vague) dynamics, the sound and meaning of the performances they describe remain beyond reconstruction. (this is especially true in the context of these African musical traditions which, as the site explains, were likely based largely around open-ended improvisation).
To attempt to recreate and draw historical conclusions from these transcriptions, then, is very difficult. But this source is much too valuable to be relegated to the remote corners of some archive, to be viewed only by a few academics, just because any recorded performance of its contents may (as critics could argue) provide a misleadingly evocative sense of ‘what really was.’ By recording and publicizing performances of these transcriptions, the contributors to Musical Passage have effectively breathed life back into the music, and the form of a digital history website offers a unique platform with which to present these performances to a wider public.

Obviously a lot of subjective interpretation is inherent to any performance of these centuries-old pieces of music, but the scholars and musicians involved with this project have managed to navigate these interpretive difficulties in multiple ways. First, as they make clear, they have drawn upon much previous research on African diasporic music to better be able to situate the pieces in their cultural context. Their choices of what instruments to perform the pieces on are influenced by both this research, and by drawings of instruments found in the same source as the musical transcriptions. (these instruments most prominently include what appears to be an early form of the banjo, which is the instrument featured in most of the site’s recordings).

screen-shot-2017-01-26-at-11-01-10-pm
Instruments depicted in Sloane’s Voyage to the Islands (1707)

Second, the site offers an open-ended approach that engages the historical imagination of the user while avoiding the potential pitfalls of this sort of project. The site presents multiple recordings of each piece. Each piece has one version played by a solo instrument (with accompanying vocals where applicable), which follows the transcription fairly closely. Each then has a second version which elaborates on the transcriptions to better convey the context in which the pieces were played. Most add other instruments (most crucially percussion), to reflect that most of this music was likely played by groups of musicians. The piece “Papa,” meanwhile, instead offers an extended version of the piece played on a different instrument, in a bid to better convey the likely open-ended and ‘circular’ nature of the piece, while also underlining that we don’t actually know what sort of instrument was used to play it.  As the site stresses,

Clearly, we cannot today capture exactly how these songs sounded then. These recordings are [an] invitation to listen and reflect on what you hear.

Offering multiple interpretations of these pieces, then, helps to avoid presenting a misleading or definitive sense of what they really sounded like, emphasizing that there is much we simply don’t know about the nature of these performances. At the same time, carefully elaborating on the transcriptions by adding other instruments, or by extending the arrangement, contributes to the listener’s understanding of the context in which these pieces were played. The ambient backdrop of crashing waves added in the site’s intro page also further engages the historical imagination of the listener, inviting them to imagine what it may have been like to attend one of these musical gatherings by situating the music in a physical environment.

Overall, Musical Passage’s open-ended approach and intuitive design is very engaging. It educates about the early history of Afro-Caribbean music by breathing life back into 300-year-old musical transcriptions, while also being careful to make clear the interpretive difficulties of doing so. The platform of a digital history website allows the dissemination of these recorded performances among a wide public audience, to an extent that would not be matched by a traditional article publication or a local performance. Moreover, the site’s creators also openly encourage more people to record and share their own performances or reinterpretations of the music. This speaks to one of the project’s most fundamental aims: to keep these rare pieces of music alive. Therefore, as a piece of music history, Musical Passage is an invaluable resource.