Tag: LDA

  • The Story of Stopwords: Topic Modeling an Ekphrastic Tradition

    The Story of Stopwords: Topic Modeling an Ekphrastic Tradition

    The following paper was first presented on July 9, 2014 at DH2014 in Lausanne, Switzerland. The slides can be found here.

    The story I’d like to tell you today is about exploring hermeneutic spaces between poetic convention and artistic invention. It’s a story about the strict algorithmic logic of topic models confronted by the rich ambiguity of poetry to revise long-standing critical assumptions about the syntactic and semantic compression of poetry. It might even be considered a cautionary tale about how exclusively relying on close reading of texts diminishes our interpretive reach. The story I’m going to tell you features the critical tradition of ekphrasis—poetry to, for, and about the visual arts—as it unfolds through a digitally-enabled, critical deformance that occurs while preparing and modeling a corpus of 4,500 English-language poems with the latent dirichlet allocation algorithm, commonly called topic modeling. [Slide 2]

    The purpose of my story is three-fold: First, to demonstrate through a case study the influence of stopword removal on topic models of poetic corpora. Second, to suggest one way we may respond to the concerns of skeptical poetry scholars about the practical and methodological constraints of distant reading practices such as topic modeling. And, third, to reveal one way in which my understanding of ekphrasis has transformed through an active engagement with the formal constraints of distant reading practices. [Slide 3]

    The work I’m going to discuss today is part of a larger project in which I look at the critical and poetic tradition of ekphrasis, its genre definitions and conventions, as well as its treatment of women as active participants in a tradition of looking at and writing about the visual arts. In the twentieth and twenty first centuries, poetic engagements with the visual arts, in this genre called ekphrasis—have drawn from and reshaped a Western tradition of viewing, describing, creating, and narrating images. From W.H. Auden’s “Musée des Beaux Arts” to Jorie Graham’s “San Sepolcro,” poetic conversations between the visual and verbal arts are as active as they have ever been since Homer’s first description of the shield of Achilles in The Iliad (18.483-601). Likewise over the 20th century, our critical understanding of the genre has also evolved. For example, Jean Hagstrum’s 1950 book The Sister Arts describes ekphrasis as “pictoral” poetry, while New Critics Joseph Frank and Murray Kreiger adopt the term to describe the realization of an image through poetic form. More recently W.J.T. Mitchell in 1992 redefines ekphrasis as “the verbal representation of visual representation”—a move that signals a methodological shift away from a metaphorical comparison of the arts of poetry and painting toward a semiotic and cultural approach.

    Similarly, my own study addresses methodological challenges faced by scholars of ekphrasis as they consider its active tradition. As significant and transformative as Mitchell’s contributions have been, “Ekphrasis and the Other” is limited in ways that are difficult to ignore in today’s critical landscape. Specifically, Mitchell’s “Ekphrasis and the Other” is based on four poems by white, male either British Romantic or American Modernist authors:[Slide 4]

    • Wallace Stevens’ “Anecdote of a Jar”;
    • William Carlos Williams’ “Portrait of a Lady”;
    • John Keats’ “Ode on a Grecian Urn”;
    • Percy B. Shelley’s “On the Medusa of Leonardo Da Vinci in the Florentine Gallery.”

    In fact, whether it be in essay or monograph form, the most poems considered in any critical definition of the genre is 75; and of all of those collections, only a handful have included examples of ekphrastic poetry by women. So when Mitchell’s essay confidently concludes: [Slide 5]

    My examples are canonical in their staging of ekphrasis as a suturing of dominant gender stereotypes into the semiotic structure of the imagetext, the image identified as feminine, the speaking / seeing subject of the text identified as masculine.

    It’s difficult to read such a conclusion without some skepticism. While Mitchell’s and others such as James A.W. Heffernan suggest that the “treatment of the ekphrastic image as female other is commonplace in the genre,” numerous feminist scholars, such as Beth Loizeaux, Barbara Fischer, Jane Hedley, Anne Keefe, and Bonnie Costello caution us against the dangers of reading ekphrasis as an overly-determined gendered contest, suggesting that women in the 20th century write with a much more nuanced attitude toward the gendered dynamics of ekphrasis. However, moving beyond Mitchell’s description has proven difficult. The genre’s saturation with frequently used and gendered terms such as “see,” “say,” “look” and “still” ” – all associated with either a dominant male gaze or a still, passive female figure—complicate our access to other critical approaches at close distance.

    I came to this particular project, then, with an explicit intent to uncover methodological means for expanding the number of poems we can consider at one time, as well as a method for detecting latent patterns across texts that were perhaps to this point obscured from view. My project, Revising Ekphrasis, asks what opportunities distant reading practices, such as topic modeling, present scholars of ekphrasis as a means for casting the widest possible net to detect latent patterns across an increasingly diverse corpus of ekphastic poetry. [Slide 6] The Revising Ekphrasis dataset includes 4,771 poems (documents) from three kinds of poetic corpora: The American Academy of Poets, Poetry Daily, and print anthologies of ekphrastic poetry. The collection, breaks down in the following way: 2,466 from the American Academy of Poets website; 373 from Poetry Daily (2011 – 2012); 34 from a print anthology titled The Gazer’s Spirit by John Hollander; 79 poems discovered through Poets on Paintings: A Bibliography by Robert Denhem; 19 poems by Jorie Graham, Carol Snow, Barbara Guest, and Cole Swenson. I also lightly described the dataset as “ekphrastic,” “nonekphrastic” or “unknown” and discovered that about 425 could be identified as ekphrastic.

    I chose topic modeling because critical treatments of ekphrasis often hinge upon perceptions that the genre uses a higher ratio of words relating to “stillness” and words relating to “look” and “see” than other genres. Questions about recurring patterns of term across similar texts are excellent kinds of questions to explore with topic modeling in theory; however, topic modeling best practices suggest that words such as “see” and “say” and “still,” which are considered high-frequency, low-semantic value terms known as stopwords. [Slide 7] Stopwords, as you likely already know, are typically removed from the dataset entirely in order to improve the quality of the model’s results. So, I was confronted right away with a dilemma, could I still get useful, reliable, and interesting results from topic modeling if all the words considered to be most influential to the genre’s definition were removed?

    As a literary critic, the idea of removing so many words from a dataset of poetry seemed preposterous at first. Auden’s opening line, for example, “About suffering they were never wrong, / the Old Masters: how well they understood…” would be dramatically different if instead it read:[Slide 8] “suffering wrong masters understood.” Robert Browning’s pivotal line in “My Last Duchess,” [Slide 9]”she looked on, and her looks went everywhere” justifying the duke’s “silencing” of his wife would be removed [Slide 10] entirely from the text of the poem if I were to use the default preprocessing steps for MALLET, the software I chose to use for my topic modeling.

    In order to feel any sense of confidence about the kind of results I would be generating, I first needed to know how the “standard” stopword removal process would influence results, so I designed a test to see how the presence of stopwords affected the usefulness of the topic keyword distributions the LDA produced.  [Slide 11] Without introducing the LDA process in detail here, what is useful to know is that LDA is an algorithm that sorts through documents and creates groupings of words that are most likely to co-occur in similar documents. Topic modeling produces distributions of words, vocabularies, that are most likely to be found in similar documents, which are called “topics,” and these distributions frequently enjoy thematic coherence.

    My tests were designed to foreground the presence or absence of important ekphrastic words, such as look, see, still, and stillness in order to expose how their absence or presence influenced the topic model’s keyword distributions—those lists of key words that topic models spit out. I imported the entire text dataset into MALLET using four different stopword lists, then ran a 40 topic model. [Slide 12]

    In the first test, I skipped preprocessing altogether, leaving every word in tact in the corpus. In the second, I heavily edited the MALLET stoplist, so that only 200 words were removed from the dataset, including adverbs (accordingly, actually, and usually); articles (a, the); and conjunctions (and, or). The third test only slightly modifies the stoplist, removing 510 high-frequency words, but leaving words frequently associated with ekphrasis in the text to be modeled, such as: after, before, look, looking, looked, looks, said, saw, say, saying, says, see, sees, seeing. Finally, I ran the last test using MALLET’s default stoplist of all 538 words.

    Rather than focusing on the significance of the individual topics to the composition of the whole model, I want to focus on the composition of the topics themselves. For those unfamiliar with reading topic model results, the number on the far left represents the topic number, [Slide 13] in this case from 0-39 because 40 topics were requested. The next number, called a hyperparameter estimation, shows the model’s prediction as to how much of the collection might be described by each topic. Next, to the right of the hyperparameter estimation are the top 20 words associated with the topic in descending order from most likely to least likely called “keys”.

    When too many high-frequency words were left in the dataset, the signal-to-noise ratio became too high to interpret the word distributions. Results from the first test are almost too difficult to read. While literary scholars of ekphrasis have understood words such as “see” and “say” and “look” to be more prominent in ekphrastic poetry than in other poetic genre, the degree to which those words are more frequently used is not significant enough to create an “ekphrastic” topic. For example, [Slide 14] topic 8 is dominated by frequently used verbs: “was, and, a, had, were, that, in, it, to, said, but, could, did, came, when, saw, they, then, I, one.” While the topic includes the word “saw,” which might be interesting in terms of understanding ekphrastic poetry, in this case it does little to identify a trend regarding ekphrastic “looking.” Instead, the topics with the highest proportions, such as Topic 1 [Slide 15], represent distributions of articles and pronouns, offering little insight into the texts themselves.

    The results from the second experiment with about half the words from the MALLET stoplist removed were not much better, as the model again appears heavily influenced by pronouns and prepositions. [Slide 16] For example, Topic 2 forms around collective identities with key words including: they, their, them, are, in, by, men, children, on, have, women, up, themselves, see. Similarly, [Slide 17] Topic 22 combines pronouns with collective bodies, or many body parts, such as: our, us, ourselves, from, how, together, live, bodies, even, heads. Furthermore, despite what we might expect from the art-oriented vocabulary in topic 33 [Slide 18], which includes words such as “by, new, fear, modern, art, times, painting, museum, mr, order, model, artist, calm….”, none of the ekphrastic poems in the dataset are expected to draw more than 4% of their language from that topic. Parsing the exact relationship between the documents that draw heavily from Topic 33 is not a matter of nuance. Instead the group seems to be primarily created around the use of the word “by.” Having specifically included “by” in the model does seem to have made a difference in terms of locating and identifying poems in which “by” accompanies other kinds of words, many of which relate to other visual aesthetic objects, but sorting through the topic is about as useful as conducting any kind of close reading of the word “by” in a poetic collection. Consequently, the results are too disperse and don’t help us to answer the questions we hoped to ask about “stillness” or “looking.”

    Perhaps the most demonstrative and telling difference between the results in the third and fourth test results is that topics where the words “look,” “see,” “still,” “at,” and “before” are included the key term distributions, as they are in tests 3 and 4, where the concentration of ekphrastic poems in a single topic is more likely when the stopwords are removed than when they are left in.[1] In other words, in the third test, when words such as “see” “say” and “still” are left in, ekphrastic poems are less likely to draw from the same topic than in the fourth test where the words are removed. For example, in the third test, 88 ekphrastic poems were predicted to draw more than 1% of their language from Topic 13. [Slide 19] Topic 13, likewise, is the most dominant topic across the collection and is associated with a predicted 72% of the poems overall, meaning the words with the greatest weight in Topic 13—“night, at, light, dark, sun, sky, day, wind, sleep, still”—are likely to be found in at least 72% of the corpus as a whole. Other topics from which many ekphrastic poems draw more than 1% of their language include topics [Slide 20] 13, 15, 16, 19, and 29. In the end, by returning to the ekphrastic poems in the dataset and sifting through the topics they are most likely associated with, what becomes evident is that the presence of “see” and “saw” or “look” and “looked” in the dataset tends to decrease the likelihood that the topic model will cluster ekphrastic poems together.

    Unexpectedly to me, leaving words such as “still” and “see” in the dataset actually impeded topic cohesion among ekphrastic poems, detecting affinities between texts that had more to do with the tense of the word than with its semantic function. The fourth test, however, which uses MALLET’s default stoplist, yielded the most salient and viable results. For example, Topic 35 predicts that 7% of the collection includes language that draws heavily from the visual arts: color, stop, painting, art, artist, painted, model, perfect, cover, museum, wrong makes, reason, wanted, witness, change, completely, case, light, hope. The model’s prediction closely reflects our pre-existing knowledge from the metadata that approximately 5% of the database’s poems are ekphrastic. These findings are doubly relevant. First, by so [Slide 21] closely estimating the number of poems that draw from the language of visual art the model promises a higher likelihood of identifying the distributions of language my study wants to explore.

    Second, [Slide 22] by estimating slightly more texts draw from language closely allied to the visual arts, the fourth test offers the tantalizing possibility of discovery. Topic 11, which describes 64% of the entire corpus but is not the most heavily weighted topic in the model, contains 125 ekphrastic poems—almost 50% of the poems with the category tag, “ekphrastic.” Looking at the key terms associated with the topic 11, one might describe it as the “body part / physical feature” topic (eyes, back, body, face, hands, hand, head, arms, feet), (open, air, inside, world, small, close), and shade (light, dark, white, black). While it’s not inconceivable that ekphrastic poems draw heavily from the language of physical description, it seems interesting the degree to which this is the case. Interestingly, too, while most of the poems most closely associated with Topic 11 are ekphrastic, the poem with the highest estimated proportion of language from it is William Carlos Williams’s poem “Danse Russe”—which I had originally tagged “nonekphrastic.” Williams, however, wrote entire volumes of ekphrastic poetry, of which Pictures from Breughal is an example. The Ballets Russes, which transformed 20th century ballet, combined efforts across the fine arts. Visual artists, such as Pablo Picasso, Henri Matisse, and Juan Gris collaborated with Russian and French choreographers, producing sets, costumes, curtains, posters, and even programs for performances.[2] While there is no textual, or as far as I can tell critical, discussion of this particular poem in terms of the visual arts Williams’s prolific ekphrastic writing, his close relationship to visual artists, and the shared language between “Danse Russe” and more than half the other ekphrastic poems in the collection present a rich opportunity for further exploration, particularly in light of some critics’ assertion that Williams’ representation of female bodies differs significantly from his male contemporaries.

    Herein, lies the critical point that I’d like to leave you with today. It is the fourth test, the one in which the language seen as “conventionally ekphrastic” is removed that opens up future hermeneutic possibility. Akin to Stephen Ramsay’s suggestion in Reading Machines that algorithmic criticism is most useful to the literary scholar when it exposes the text to new hermeneutic potential, topic models’ are most useful to the study of poetry, and in this case ekphrasis, when the method exposes our scholarly assumptions through de-familiarization: [Slide 23]

    Literary-critical insight begins with a change of vision—what Wittgenstein called the “drawing of an aspect.” (Philsophical 194). Sometimes… the noticing is the result of some sort of overt manipulation of the text. We read out of order, we translate and paraphrase, we look only at certain words or certain constellations of surrounding the text. The text hasn’t changed its graphic content any more than the duck-rabbit changes between one’s seeing it one way one moment and another the next. But the text quite literally assumes a different organization from what it had before. Once a new aspect/pattern has been discovered, one immediately begins to test the viability of that pattern (47-8).

    In the case of these topic model distant readings, the radical extrusion of words considered vital to the identification of a poetic genre exposes rich opportunities to explore latent patterns and connections heretofore obscured by the over-presence of a limited vocabulary. Removing stopwords, a form of textual “deformance” akin to Jerome McGann and Lisa Samuel’s use of the term, leads us to wonder what more be gained by adjusting the aperture of our scholarly lens to reconsider the ekphrastic tradition at scale. The value of algorithmic criticism, brought to bear on the Revising Ekphrasis corpus, leverages what Ramsay describes as “a desire to use the narrowing forces of constraint to enable the liberating visions of potentiality” (32). Much in the vein of Matthew Jockers, Ted Underwood, Andrew Goldstone, Ben Schmidt, and Lauren Klein’s work with corpora of fiction, literary genres, literary criticism, and archival materials what the story of stopwords and the tradition of ekphrasis suggests is that a wealth of hermeneutic potential exists for in topic modeling for scholars of poetry, as well. [Slide 25]

    [1] The exact list of words included in test 2 that were not included in test 1 can be found in Appendix B.

    [2] “Visual Art and the Ballets Russes.” Ballet Russes Cultural Partnership. Boston University. Web. 16 Sept. 2012.

    Denham, Robert D. Poets on Paintings: A Bibliography. Jefferson, N.C.: McFarland, 2010. Print.

    Frank, Joseph. The Idea of Spatial Form. New Brunswick: Rutgers University Press, 1991. Print.

    Hagstrum, Jean H. The Sister Arts: The Tradition of Literary Pictorialism and English Poetry from Dryden to Gray. Chicago: University of Chicago Press, 1987. Print.

    Hedley, Jane, Nick Halpern, and Willard Spiegelman. In the Frame: Women’s Ekphrastic Poetry from Marianne Moore to Susan Wheeler. Newark, DE: University of Delaware Press, 2009. Print.

    Heffernan, James A. W. Museum of Words: The Poetics of Ekphrasis from Homer to Ashbery. Chicago: University Of Chicago Press, 2004. Print.

    Hollander, John, and Various Authors. The Gazer’s Spirit: Poems Speaking to Silent Works of Art. 1st ed. Chicago: University Of Chicago Press, 1995. Print.

    Krieger, Professor Murray. Ekphrasis: The Illusion of the Natural Sign. Baltimore: The Johns Hopkins University Press, 1992. Print.

    Loizeaux, Elizabeth Bergmann. Twentieth-Century Poetry and the Visual Arts. 1st ed. Cambridge, UK; New York: Cambridge University Press, 2008. Print.

    McCallum, Andrew Kachites. MALLET: A Machine Learning for Language Toolkit. N.p., 2002. Web.

    McGann, Jerome. Radiant Textuality: Literature After the World Wide Web. New York, NY: Palgrave Macmillan, 2004. Print.

    Mitchell, W. J. T. Iconology: Image, Text, Ideology. Chicago: University of Chicago Press, 1986. Print.

    Mitchell, W. J. T. Picture Theory: Essays on Verbal and Visual Representation. Chicago: University of Chicago Press, 1995. Print.

    Ramsay, Stephen. Reading Machines: Toward an Algorithmic Criticism. 1st Edition edition. Urbana: University of Illinois Press, 2011. Print.

    Scott, Grant F. “The Rhetoric of Dilation: Ekphrasis and Ideology.” Word & Image 7.4 (1991): 301–310. Taylor and Francis+NEJM. Web. 12 Oct. 2012.

  • Teaching LDA with the Topic Modeling Game

    Teaching LDA with the Topic Modeling Game

    The following post was featured as a Digital Humanities Now Editor’s Choice entry on April 2, 2013.

    In February, I visited Matthew Kirschenbaum’s #ENGL668K Introduction to Digital Humanities course at the University of Maryland, and I brought to class an activity that I had been mulling over in my own mind for a long time, called the Topic Modeling Game.  The game is designed to teach the basic principles of topic modeling with LDA through engaged, constructivist, and problem-based techniques.

    As I was learning about LDA myself, I realized that I was essentially playing this game in my head over and over again: following through how I thought LDA worked, learning where I made mistakes, revising my assumptions, and playing the game all over again.  When I went to write about the results of my topic modeling experiments for the Revising Ekphrasis project in my dissertation, what I discovered is that I really needed a way to explain topic modeling such that someone who had no knowledge of the methodology could read the results of my experiments and trust my conclusions.

    That process led to a written explanation of LDA that will appear in a future article in the Journal of Digital Humanities.  In the essay, I create a hypothetical situation to explain the assumptions LDA makes about natural language texts in order to produce its results.  The example walks readers through the process of figuring out what produce is available at a farmers’ market that they have never been to themselves and asks the reader to consider the problem from a quantitative perspective.  When I created that explanation, I did it by playing this game in my head.  So, I thought, perhaps this could be an effective way of teaching LDA, as well.

    After tweeting something about the Topic Modeling Game, other DH instructors requested copies of the game.  I absolutely wanted to share, but I also wanted to learn from other teachers’ experiences.  More importantly, I wanted other new teachers to benefit from the experience of those who had already tried it.

    Meanwhile, there’s been a lot of conversation about the value of public, open repositories of data.  ProfHacker has had several recent posts about using GitHub to revise documents (see Getting Started with a GitHub Repository and Forks and Pull Requests in GitHub).  Also, Matt Burton presented a helpful introduction to Git at MLA 2013 in the Scaling and Sharing: Data Management in the Humanities special session #s586.

    What better way to share a lesson plan with peers, I figured, than to create a repository for it on GitHub, to invite others to use it and to ask that they share their results and add their changes and revisions back to the repository?

    As a result, I created a GitHub public repository for the Topic Modeling Game.  Currently, in the repository, there are two Word documents.  One includes instructions and background information for teachers.  The other is a rudimentary hand-out to use to begin the game as an in-class (face-to-face) assignment.  The instructor document includes a list of materials that could be used for the lesson, but I have not uploaded sheets of “sample words” to use—at least not yet.

    There is plenty of room to edit, improve, revise, innovate, and share.  The one thing that I do ask is that if you download and use the Topic Modeling Game, you contribute to the repository by adding lessons you have learned, revisions you made, and suggestions for improvement.

    So, bring your forks (again, you may want to read Konrad Lawson’s recent post on forking and pulling with git if you’re unfamiliar with the process) and dig on in.  I’m looking forward to hearing back from those who try it.

  • Some Assembly Required: Understanding and Interpreting Toics in LDA Models of Figurative Language

    Some Assembly Required: Understanding and Interpreting Toics in LDA Models of Figurative Language

    The following is a small part from a much larger work in progress (my dissertation) about the potential to use latent Dirichlet allocation (LDA) to do exploratory work with highly figurative language.  In fact, my project uses LDA to model various iterations of an approximately 4,500 poem dataset (the majority of which are from the 20th century), and to consider the composition of that dataset in relationships to a smaller subset of the data that could be described as belonging to a poetic tradition called ekphrasis: poetry to, for, and about the visual arts.  There’s no way that I could begin to get into all the nitty gritty details about the rest of the project in this one blog post, so with apologies, I’m going to begin in media res, assuming that you know this much: probabilistic topic modeling, and in particular LDA, is a way of looking for patterns in large collections of text.  In previous posts I’ve mentioned that there are many good posts on what LDA is and how some humanists are using it.  Most recently Scott Weingart produced a blog post called “Topic Modeling for Humanists: A Guided Tour” that adds to the much needed collection of “How to get started” conversations.  Rather than focusing on how topic modeling can be useful to you, this is a post about how you, dear Reader, need to read our results—at least those of us who are working with figurative texts and particularly those of us working with poetry, the most figurative of them all.

    If you’re just getting started, it’s important to begin with the following knowledge: data mining in any form makes two assumptions that Ian H. Witten, Eibe Frank, and Mark Hall point out in their introduction to the topic and to their graphical interface data mining software Weka.  They remind us data mining results need to be actionable and they need to be comprehensible.  I’ll go into what I think that means for my work, but suffice it to say, topic modeling assumes that texts, though amorphous, don’t hide information.  In fact, text mining in general assumes that writers go to great lengths to make clear, unambiguous arguments.  Computer scientists make that assumption because LDA was written to deal with large repositories and collections of non-fiction text.  When you’re reading the journal Science, for example, you don’t see lines like:

    Little lion face

    I stooped to pick
    among the mass of thick
    succulent blooms, the twice
    streaked flanges of your silk

    sunwheel relaxed in wide
    dilation, I brought inside,
    placed in a vase. Milk
    of your shaggy stem

    sticky on my fingers, and
    your barbs hooked to my hand,
    sudden stings from them
    were sweet…

    May Swenson wasn’t writing for Science, and her poem is about more than dandelions.  Pretty much any human reading that poem, even my undergraduates, get that this is a poem about sex.  Science’s editors would never publish this; however, they may publish and have published plenty of articles about sex, reproduction, and the propagation of flora and fauna.  The terms they use, though, strive against ambiguity, while poetry revels in it.  We don’t have a well-established way of interpreting topics that account for poetry’s lush ambiguity, but we need to because it would be a mistake to read a topic with the keywords: wind, sky, light, trees, blue, white, snow… when generated from a collection of poems the same way you would read and understand it in, say, David Blei’s 100-topic model of Science.

    Rather than reposting Blei’s images here, I suggest that readers interested in understanding LDA look at his article from Communications of the ACM, because his illustrative examples on the first and second pages do an excellent job of showing how topics are generated.   Blei’s results are these wonderfully identifiable topics that make such sense: of course, we can interpret topic as the genetics topic because it is comprised of words like gene and dna and another as the evolutionary biology topic because it is made up of words like survival and mutation.

    So while the classic examples of topic models produce semantically and thematically coherent keyword distributions, should we expect highly figurative texts, particularly poems but not exclusive of other forms of highly figurative texts such as fiction and drama, to form around the same kind of thematic topics?  Returning once again to Blei’s most accessible article for humanists, he writes: “The interpretable topic distributions arise by computing the hidden structure that likely generated the observed collection of documents.” Blei clarifies his statement in a footnote which reads: “Indeed calling these models “topic models” is retrospective—the topics that emerge from the inference algorithm are interpretable for almost any collection that is analyzed.  The fact that these look like topics has to do with the statistical structure of observed language and how it interacts with the specific probabilistic assumptions of LDA” (Blei “Introduction” 79).  In other words, the topics from Science scan as comprehensible, cohesive topics because the texts from which they were derived strive to use language that identifies very literally with its subject.  The algorithm, however, does not know the difference between texts that tend to be more literal than figurative.  The same process for identifying topics applies for both literal and figurative texts: topics are a distribution over a fixed vocabulary.  The first stage of a topic modeling experiment with poetry, then, is a matter of determining what those distributions look like and whether or not they can be useful.

    What would the same illustrative example Blei created for Science look like in an LDA model run on a corpus of poetry?  The poem below translates the LDA intuitions described in Blei’s article to the situation of a poem in a dataset of 4,500 poems.  For copyright purposes, I’m going to remove from the poem the words that would be removed during preprocessing (the stopwords) of Anne Sexton’s “The Starry Night” but if you want to see the whole poem, look here.

    Anne Sexton’s Starry Night with stopwords removed.

    In the case of Anne Sexton’s “Starry Night,” LDA assumes that the three most prominent topics in the poem are 32, 2, and 54.  In the chart below, I list the topic assignment at the top with estimated distribution of the topic across the document.  Under each topic is a list of the top 15 keywords most strongly associated with those topics.

    Topic 32 (29%)Topic 2 (12%)Topic 54 (9%)
    night
    light
    moon
    stars
    day
    dark
    sun
    sleep
    sky
    wind
    time
    eyes
    star
    darkness
    bright
    death
    life
    heart
    dead
    long
    world
    blood
    earth
    man
    soul
    men
    face
    day
    pain
    die
    tree
    green
    summer
    flowers
    grass
    trees
    flower
    spring
    leaves
    sun
    fruit
    garden
    winter
    leaf
    apple

    LDA analysis, then, reads Anne Sexton’s “Starry Night” as containing 25% of its words from topic 32, which seems generally to draw on language associated with time of day, 12% of its language from topic 2 which includes many words about death and dying, and 9% of its language from the natural environment.  Strong coherence among keywords in topics 32 and 54 simplify the interpretive task of assigning labels to them; however, topic 2 is not so easily labeled.  The terms “death, life, heart, dead, long, world” are extremely broad, and to my mind easily misread or misinterpreted without the context of the data that it describes.  Only in light of referring back to “The Starry Night” (and other poems closely associated with topic 2) can we develop a sense of hermeneutic confidence about the comprehensibility of such results, which are discussed further on.

    I’m doing a lot of cutting from the original document in which I make this argument (insert shameless plug here regarding my dissertation), so please forgive some of the logical leaps.  I want to jump ahead to the point: Why do we care about what kinds of topics these are and how does the relate to the need for close readings?

    Essentially, in my dataset, I found four kinds of topics are most likely to appear when topic modeling poetry: OCR and foreign language topics; “large chunk” topics (a document larger than most of the rest with language that dominates a particular topic); semantically evident topics; and semantically opaque topics.  I’ll describe the latter two here:

    1.)    Semantically evident topics—Some topics do appear just as one might expect them to in the 100-topic distribution of Science in Blei’s paper.  Topics 32 and 54 illustrated above in Anne Sexton’s “Starry Night” exemplify how LDA groups terms in ways that appear upon first blush to be thematic, as well.  Our understanding, though, of these semantically evident topics as they are generated by highly figurative texts requires a bit of refinement.  It may be accurate to say that time of day and natural landscapes are topics in “Starry Night.”  After all, Sexton does describe a painted landscape under the stars, but it would not be correct to say that 29% of the document is “about” the time of day.  As literary scholars, we understand that Sexton’s use of the tumultuous night sky depicted by Vincent Van Gogh provides a conceit for the more significant thematic exploration of two artists’ struggle with mental illness.  Therefore, it is important not to be seduced by the seeming transparency of semantically evident topics.  These topics reflect most powerfully Ted Underwood’s definition of “LDA topic” as “discourse.”  In other words, topics form around a manner of speech, and the significant questions to be asked regarding such topics have to do with what we learn about the relationships between forms of discourse associated with particular topics across documents within a specific dataset.

    2.)    Semantically opaque topics—Some topics, such as topic 2 in the “Starry Night” example are not immediately apparent.  In fact, I found them to be discouraging the first time I started running LDA models of the dataset because they are so difficult to synthesize into the single phrases used by so many of the researchers in not only computer sciences but digital humanities, as well.  Determining a pithy label for a topic with the keywords, “death, life, heart, dead, long, world, blood, earth…” is virtually impossible until you return to the data, read the poems most closely associated with the topic, and infer the commonalities among them:

    TopicProportion     Title
    20.535248643When to the sessions of sweet silent thought (Sonnet 30)
    20.533343438By ways remote and distant waters sped (101)
    20.517398877A Psalm of Life
    20.481152152We Wear the Mask
    20.477938906The times are nightfall, look, their light grows less
    20.472091675The Slave’s Complaint
    20.451175606The Guitar
    20.447100571Tears in Sleep
    20.446314271The Man with the Hoe
    20.437962153A Short Testament
    20.433767746Beyond the Years
    20.433152279Dead Fires
    20.429638773O Little Root of a Dream
    20.427326132Bangladesh II
    20.425835136Vitae Summa Brevis Spem Nos Vetat Incohare Longam

    Topic 2 is interesting for a number of reasons, not the least of which is that even though Paul Laurence Dunbar’s “We Wear the Mask” never once mentions the word “death,” the language Dunbar uses to describe the erasure of identity and the shackles of racial injustice are identified as drawing heavily from language associated with death, loss, and internal turmoil—language which “Starry Night” indisputably also draws from.  To say that this is a topic about “death, loss, and internal turmoil” is overly simplistic.  Just as semantically evident topics require interpretation, so do semantically opaque topics.  While the former tends to center around images, metaphors, and particular literary devices, the latter topic often emphasizes tone.  Words like “death, life, heart, dead, long, world” out of context tell us nothing about an author’s attitude or thematic affinities between poems, but when a close reader scales down into the compressed language of the poems themselves that draw from the topic’s language distribution, there are rich deposits of hermeneutic possibility.  There’s a lot that could be said about elegy here and the relationships between elegy and other poetic genres… but I’ll save that for another post.

    At long last, the point: if we assume that the “semantically evident” topics are actually about the words by themselves, we’re missing something important.  Semantically evident and semantically opaque topics in LDA models of highly figurative texts must be the starting point for an interpretive process.  It is incumbent upon us as digital humanists who use this methodology to explain that a topic with keywords like “night, light, moon, stars, day” isn’t just about time of day.  More likely, it’s about the use of time of day as images, metaphors, and other figurative proxies for another conversation and none of that is evident without a combination of close and “networked” reading.  These four topic types appear in every model to varying degrees based on the number of topics input during the construction of my LDA models and represent the difference between topic models of figurative language as opposed to topic models of non-fiction, journalistic, or academic prose.  As a result, reading, navigating, and interpreting topics in a figurative dataset requires a slightly different approach than reading, navigating, and interpreting models of other kinds of text collections.  Moreover, understanding topics requires a networked interpretive strategy.  Texts need to be read in relationship to other texts in the corpus, and how that happens, what I suggest for the best practices for doing networked readings is a point I’ll have to make in the next post.

  • Why use visualizations to study poetry?

    Why use visualizations to study poetry?

    This post was a DHNow Editor’s Choice selection on May 1, 2012.

    The research I am doing presently uses visualizations to show latent patterns that may be detected in a set of poems using computational tools, such as topic modeling.  In particular, I’m looking at poetry that takes visual art as its subject, a genre called ekphrasis, in an attempt to distinguish the types of language poets tend to invoke when creating a verbal art that responds to a visual one.  Studying words’ relationships to images and then creating more images to represent those patterns calls to mind a longstanding contest between modes of representation—which one represents information “better”?  Since my research is dedicated to revealing the potential for collaborative and kindred relationships between modes of representation historically seen in competition with one another, using images to further demonstrate patterns of language might be seen as counter-productive.  Why use images to make literary arguments? Do images tell us something “new” that words cannot?

    Without answering that question, I’d like instead to present an instance of when using images (visualizations of data) to “see” language led to an improved understanding of the kinds of questions we might ask and the types of answers we might want to look for that wouldn’t have been possible had we not seen them differently—through graphical array.

    Currently, I’m using a tool called MALLET to create a model of the possible “topics” found in a set of 276 ekphrastic poems.  There are already several excellent explanations of what topic modeling is and how it works (many thanks to Matt Jockers, Ted Underwood, and Scott Weingart who posted these explanations with humanists in mind), so I’m not going to spend time explaining what the tool does here; however, I will say that working with a set of 276 poems is atypical.  Topic modeling was designed to work on millions of words, and 276 poems doesn’t even come close; however, part of the project has been to determine a threshold at which we can get meaningful results from a small dataset.  So, this particular experiment is playing with the lower thresholds of the tool’s usefulness.

    When you run a topic model (train-topics) in MALLET, you tell the program how many topics to create, and when the model runs, it can output a variety of results.  As part of the tinkering process, I’ve been working with the number of topics to have MALLET use in order to generate the model, and was just about to despair that the real tests I wanted to run wouldn’t be possible at 276 poems.  Perhaps it was just too few poems to find recognizable patterns.  For each topic assignment, MALLET assigns an ID number to the topic and “topic keys” as keywords for that topic.  Usually, when the topic model is working, the results are “readable” because they represent similar language.  MALLET would not call a topic “Sea,” for example, but might instead provide the following keywords:

    blue, water, waves, sea, surface, turn, green, ship, sail, sailor, drown

    The researcher would look at those terms and think, “Oh, clearly that’s a nautical/sea/sailing” topic, and dub it as such.  My results, however, on 15 topics over 276 poems were not readable in the same way.  For example, topic 3 included the following topic keys:

    3          0.04026           with self portrait him god how made shape give thing centuries image more world dread he lands down back protest shaped dream upon will rulers lords slave gazes hoe future

    I don’t blame you if you don’t see the pattern there.  I didn’t.  Except, well, knowing some of the poems in the set pretty well, I know that it put together “Landscape with the Fall of Icarus” by W.C. Williams with “The Poem of Jacobus Sadoletus on the Statue of Laocoon” with “The New Colossus” with “The Man with the Hoe Written after Seeing the Painting by Millet.”  I could see that we had lots of kinds of gods represented, farming, and statues, but that’s only because I knew the poems.  Without topic modeling, I might put this category together as a “masters” grouping, but it’s not likely.  Rather than look for connections, I was focused on the fact that the topic keys didn’t make a strong case for their being placed together, and other categories seemed similarly opaque.  However, just to be sure that I could, in fact, visualize results of future tests, I went ahead and imported the topic associations by file.  In other words, MALLET can also produce a file that lists each topic (0-14 in this case) with each file name in the dataset and a percentage.  The percentage represents the degree to which the topic is represented inside each file.  I imported the MALLET output of topics and files associated with them into Google Fusion Tables and created a dynamic bar graph that collects file-ids along the vertical axis and along the horizontal axis can be found the degree that the given topic (in this case topic 3) is present in the file.   As I clicked through each topic’s graph, I figured I was seeing results that demonstrated MALLET’s confusion, since the dataset was so small.  But then I saw this: [Below should be a Google Visualization.  You may need to “refresh” your browser page to see it.  If you still cannot see it, a static version of the file is visible here.]

    https://web.archive.org/web/20141022105827if_/https://www.google.com/fusiontables/embedviz?&containerId=gviz_canvas&q=select+col0%2C+col4+from+3650097+&qrs=where+col0+%3E%3D+&qre=+and+col0+%3C%3D+&qe=+limit+247&viz=GVIZ&t=BAR&width=500&height=500

    If the graph’s visualization is working, when you pass your mouse over the lines in the bar graph, the ones that are higher than 0.4, then the file-id number (a random number assigned during the course of preparing the data) appears.  Each of these files begin with the same prefix: GS.  In my dataset, that means that the files with the highest representation of topic 3 in them can all be found in John Hollander’s collection The Gazer’s Spirit.  This anthology is considered to be one of the most authoritative and diverse—beginning with classical ekphrasis all the way up to and including poems from the 1980s and 1990s.  I had expected, given the disparity in time periods, that the poems from this collection would be the most difficult to group together because the diction of the poems changes dramatically from the beginning of the volume to the end.  In other words, I would have expected the poems to blend with the other ekphrastic poems throughout the dataset more in terms of their similar diction than by anything else.  MALLET has no way of knowing that these files are included in the same anthology.  All of the bibliographical information about the poems has been stripped from the text being tested.  There has to be something else.  What something else might be requires another layer of interpretation.  I will need to return to the topic model to see if a similar pattern is present when I use  other numbers of topics—or if I include some non-ekphrastic poems to the set being tested—but seeing the affinity in language between the poems included in The Gazer’s Spirit in contrast to other ekphrastic poems proved useful.  Now, I’m not inclined to throw the whole test away, but instead to perform more tests to see if this pattern emerges again in other circumstances.  I’m not at square one. I’m at a square 2 that I didn’t expect.

    The visualization in the end didn’t produce “new knowledge.”  It isn’t hard to imagine that an editor would choose poems that construct a particular argument about what “best” represents a particular genre of poetry; however, if these poems did truly represent the diversity of ekphrastic verse, wouldn’t we see other poems also highly associated with a “Gazer’s Spirit topic”?  What makes these poems stand out so clearly from others of their kind?  Might their similarity mark a reason for why critics of the 90s and 2000s define the tropes, canons, and traditions of ekphrasis in a particular vein?  I’m now returning to the test and to the texts to see what answers might exist there that I and others have missed as close readers.  Could we, for instance, run an analysis that determines how closely other kinds of ekphrasis are associated with Gazer’s Spirit’s definition of ekphrasis?  Is it possible that poetry by male poets is more frequently associated with that strain of ekphrastic discourse than poetry by female poets?

    This particular visualization doesn’t make an “argument” in the way humanists are accustomed to making them.  It doesn’t necessarily produce anything wholly “new” that couldn’t have been discovered some other way; however, it did help this researcher get past a particular kind of blindness and helped me to see alternatives—to consider what has been missed along the way—and there is, and will be, something new in that.

  • Small Projects & Limited Datasets

    Small Projects & Limited Datasets

    I’ve been thinking a lot lately about the significance of small projects in an increasingly large-scale DH environment.  We seem almost inherently to know the value of “big data:” scale changes the name of the game.  Still, what about the smaller universes of projects with minimal budgets, fewer collaborators, and limited scopes, which also have large ambitions about what can be done using the digital resources we have on hand?  Rather than detracting from the import of big data projects, I, like Natalie Houston, am wondering what small projects offer the field and whether those potential outcomes are relevant and useful both in and of themselves as well as beneficial to large-scale projects, such as in fine-tuning initial results.

    My project in its current iteration involves a limited dataset of about 4500 poems and challenges rudimentary assumptions about a particular genre of poetry called ekphrasis—poems regarding the visual arts.  It is the capstone project to a dissertation in which I use the methods of social network analysis to explore socially-inscribed relationships between visual and verbal media and in which the results of my analysis are rendered visually to demonstrate the versatility and flexibility available to female poets writing ekphrastic poetry. My MITH project concludes my dissertation by demonstrating that network analysis is one way of disrupting existing paradigms for understanding the social-signification of ekphrastic poetry, but there are more methods available through computational tools such as text modeling, word frequency analysis, and classification that might also be useful.

    To this end, I’ve begun by asking three modest questions about ekphrastic poetry using a machine learning application called MALLET:

    1.) Could a computer learn to differentiate between ekphrastic poems by male and female poets?  In “Ekphrasis and the Other,” W.J.T. Mitchell argues that were we to read ekphrastic poems by women as opposed to ekphrastic poetry by men, that we might find a very different relationship between the active, speaking poetic voice and the passive, silent work of art—a dynamic which informs our primary understanding of how ekphrastic poetry operates.  Were this true and were the difference to occur within recurring topics and language use, a computer might be trained to recognize patterns more likely to co-occur in poetry by men or by women.

    2.) Will topic modeling of ekphrastic texts pick out “stillness” as one of the most common topics in the genre?  Much of the definition of ekphrasis revolves around the language of stillness: poetic texts, it has been argued, contemplate the stillness and muteness of the image with which it is engaged.  Stillness, metaphorically linked to muteness, breathlessness, and death, provides one of the most powerful rationales for an understanding how words and images relate to one another within the ut pictura poesis tradition—usually seen as an hostile encounter between rival forms of representation.  The argument to this point has been made largely on critical interpretations enacted through close readings of a limited number of texts.  Would a computer designed to recognize co-occurrences of words and assign those words to a “topic” based on the probability they would occur together also reveal a similar affiliation between stillness and death, muteness, even femininity?

    3.) Would a computer be able to ascertain stylistic and semantic differences between ekphrastic and non-ekphrastic texts and reliably classify them according to whether or not the subject of the poem is an aesthetic object or not?  We tend to believe that there are no real differences between how we describe the natural world as opposed to how we describe visual representations of the natural world.  We base this assumption on human, interpretive, close readings of  poetic texts; however, there is the potential that a computer might recognize subtle differences as statistically significant when considering hundreds of poems at a time.  If a classification program such as Mallet could reliably categorize texts according to ekphrastic and non-ekphrastic, it is possible that we have missed something along the way.

    In general, these are small questions constructed in such a way that there is a reasonable likelihood that we may get useful results.  (I purposefully choose the word results instead of answers, because none of these would be answers.  Instead the result of each study is designed to turn critics back to the texts with new questions.)  And yet, how do we distinguish between useful results and something else?  How do we know if it worked?  Lots of money is spent trying to answer this question about big data, but what about these small and mid-sized data sets?  Is there a threshold for how much data we need to be accurate and trustworthy?  Can we actually develop standards for how much data we need to ask particular kinds of humanities questions to make relevant discoveries?  In part, my project also addresses these questions, because otherwise, I can’t make convincing arguments about the humanities questions I’m asking.

    Small projects (even mid-sized projects with mid-sized datasets) offer the promise of richly encoded data that can be tested, reorganized, and applied flexibly to a variety of contexts without potentially becoming the entirety of a project director’s career.  The space between close, highly-supervised readings and distant, unsupervised analysis remains wide open as a field of study, and yet its potential value as a manageable, not wholly consuming, and reproducible option make it worth seriously considering.  What exactly can be accomplished by small and mid-scale projects is largely unknown, but it may well be that small and mid-sized projects are where many scholars will find the most satisfying and useful results.