Sunday, August 16, 2026

Lab Session1: Digital Humanities: CLiC - etc

 

Algorithms, Aesthetics, and the Shifting Mirror of Humanity: A Digital Humanities Journey

Part 1: The Poetry Debate — Can Machines Speak the Language of the Soul?

For centuries, poetry has been regarded as the ultimate bastion of human expression—the language of raw emotion and subjective experience. But what happens when an algorithm begins to write poetry? Can a computer write a sonnet that is indistinguishable from one produced by a human?

This question was historically approached through Alan Turing's legendary 1950 proposal . Turing suggested that if a machine could engage in a text-based conversation with a human with such proficiency that the human couldn't tell them apart, the machine possesses intelligence . In 2013, researchers Oscar Schwartz and Benjamin Laird applied this logic to verse, creating an online Turing test for poetry called "Bot or Not" .

The premise was simple yet unsettling: users were presented with a poem and had to guess whether its creator was human or machine . The results disrupted our comfortable assumptions. While Turing's original benchmark for passing his test was a 30% deception rate , poems on "Bot or Not" routinely fooled up to 65% of human readers .

In his illuminating TED Talk, Oscar Schwartz shared several experiments that expose the fragile boundaries of the "human" category . He contrasted an algorithmic poem derived from his Facebook feed with William Blake’s classic poem "The Fly" . While Blake was easily recognized, a matchup between the computer program Racter and poet Frank O'Hara split the audience 50/50 . The machine's cold logic had seamlessly mirrored O'Hara's vibrant, colloquial style .

Even more startling is the "reverse Turing test" involving Gertrude Stein and Ray Kurzweil’s Cybernetic Poet algorithm (RKCP) . Kurzweil's algorithm was fed Emily Dickinson’s poetry, analyzed her syntactic structures, and generated a new model . Side-by-side with Stein’s highly experimental, repetitive prose-poem, the vast majority of human judges declared that the computer's Dickinson-emulated poem was human, and that Stein's human-authored poem was a machine .

This leads to a profound existential realization: the RKCP algorithm has absolutely no semantic understanding of the words it generates . To the computer, language is merely raw mathematical material . Yet, it successfully conjures the illusion of profound human emotion. Conversely, when human poets push the boundaries of language to break conventional rules—as Stein did—their writing is often dismissed as mechanical .

Ultimately, this debate reveals that "the human" is not a cold, hard scientific fact, but an ever-shifting, culturally constructed idea . The computer operates as a mirror . If we feed it Emily Dickinson or William Blake, that is what it reflects back to us . The real question is not merely whether we can build creative machines, but what specific image of our own humanity we want to see reflected back to us in the digital looking glass.

Part 2: Testing My Human Intuition — My Experience with the Poetry Turing Tests

Armed with these philosophical insights, I decided to test my own literary intuition by taking two separate poetry Turing tests: a classroom Google Form quiz and the NPR "Human or Machine" sonnet challenge, inspired by a computational creativity competition at Dartmouth College .

My results were both humbling and highly instructive. On the Google Form quiz, I scored a moderate 6 out of 10 points . This demonstrated just how easy it is to misread the markers of human authorship in isolated, short verses. However, on the NPR sonnet challenge, I fared significantly better, scoring 5 out of 6 correct answers .

The Deception: Sonnet #2 and the Illusion of Domestic Melancholy

The single error in my NPR quiz was Sonnet #2 . Reading the lines, my gut insisted this was the work of a human poet:


The dirty rusty wooden dresser drawer. A couple million people wearing drawers, Or looking through a lonely oven door, Flowers covered under marble floors.

And lying sleeping on an open bed. And I remember having started tripping, Or any angel hanging overhead, Without another cup of coffee dripping.

I confidently selected "A Human," only to be met with a bold "WRONG!" banner . This poem was actually generated by an algorithm programmed by researchers Marjan Ghazvininejad, Xing Shi, Yejin Choi, and Kevin Knight at the University of Southern California's Information Sciences Institute .

The algorithm successfully exploited several "humanizing" literary techniques. First, it uses concrete, domestic imagery: a "dresser drawer," an "oven door," and "coffee dripping." Second, it maintains a structured, comforting ABAB rhyme scheme (drawer/door, drawers/floors, bed/overhead, tripping/dripping). Finally, the line "And I remember having started tripping" introduces a first-person pronoun ("I") linked to a subjective memory. This triply-reinforced illusion of memory, domestic reality, and formal poetic structure tricked my brain into projecting a human consciousness behind the screen.

The Successes: Deciphering the Human and the Machine

In the remaining five sonnets, my analytical training proved more resilient.

  • Sonnet #1 was correctly identified as Human . Written by Thomas Kinder, this poem possesses a genuine emotional arc . The transition from the oppressive heat of the city to the quiet, devastating grief of a hospital visit ("I walked those mornings to the hospital... to see this heat exhumes the body of that grief") has an organic emotional depth that algorithms struggle to sustain across stanzas.

  • Sonnet #3 and Sonnet #4 were also correctly identified as Human . Written by Ivy Schweitzer and Kurtis Hessel respectively, these poems feature syntactic elegance and metaphorical nuance . In Hessel's sonnet, the wordplay of "following this dress-rehearsal pain / Gives way to Joy" displays a playful, self-aware wit.



  • Sonnet #5 was correctly identified as Machine-generated . This piece was written by the "Pythonic Poet" (UC Berkeley researchers Andrea Gagliano, Emily Paul, Kyle Booten, and Marti Hearst) . While the poem successfully mimics archaic, Shakespearean vocabulary ("quake," "reproach," "boon"), its syntax is profoundly disjointed, and the transitions between lines lack semantic cohesion.

  • Sonnet #6 was another correct guess of Human . Written by Thomas Kinder, its beautiful, meditative extended metaphor of gardening felt too philosophically unified to be the product of an algorithm .

This testing experience showed me that while algorithms have mastered the formal mechanics of poetry, they often fail to create a sustained, coherent thematic progression. When a machine poem succeeds in fooling us, it is usually because we, as human readers, actively perform the emotional labor of finding meaning, stitching together disparate images into a coherent narrative of our own making.

Part 3: Corpus Linguistics in Action — The CLiC Dickens Project and "Great Expectations"

While web-based poetry quizzes offer a playful exploration of text generation, how can we study the intricate, deliberate linguistic choices of human authors on a larger scale? This is where Corpus Linguistics is invaluable. During our digital skilling curriculum, we explored the CLiC (Corpus Linguistics in Context) Dickens Project, a state-of-the-art web application developed by the Universities of Birmingham and Nottingham .

CLiC allows students to perform computer-assisted analysis of 19th-century literature . Rather than relying solely on subjective, impressionistic close readings, CLiC enables us to search massive corpora of text (including Dickens’s fifteen novels and reference corpora of other 19th-century works) to find empirical patterns, count word frequencies, analyze collocations, and isolate specific text subsets such as direct character speech ("Quotes") or narrator text ("Non-quotes") .

To see this tool in action, I undertook Activity 16: Growing up in Great Expectations, focusing closely on the thematic representations of parenting, gratitude, and moral development in Charles Dickens’s famous semi-autobiographical novel .

Activity 16.1: Being "Brought Up by Hand" in Great Expectations

In the opening chapters, the orphan protagonist Pip frequently describes how he was "brought up by hand" by his sister, Mrs. Joe Gargery . Using CLiC, I ran a concordance search for the phrase "up by hand" across the novel, retrieving exactly 14 instances.

The phrase undergoes a severe, ironic semantic shift depending on who is speaking . In the historical Victorian context, "brought up by hand" meant being fed by bottle or spoon rather than being breastfed by a mother or wet nurse. However, Mrs. Joe overlays this with physical violence . She literally uses her physical "hand" to beat Pip and Joe, often employing her wax-ended cane "Tickler."

Pip, reflecting as the adult narrator, highlights this abuse in Line 11 of the concordance:

"I had cherished a profound conviction that her bringing me up by hand, gave her no right to bring me up by jerks" 

This line is a masterpiece of Dickensian wit, contrasting the formal childcare idiom with the physical reality of being shaken and struck.

Moreover, the concordance lines reveal that the adults in Pip’s life weaponize the phrase to enforce a culture of toxic guilt and forced submission . Mr. Pumblechook instructs Pip: "...be grateful, boy, to them which brought you up by hand" (Line 5) and "be for ever grateful" (Line 12) . Even Mrs. Hubble contemplates Pip as a burden (Line 5) . The concordance shows that the phrase is never used to convey warmth or maternal care; rather, it is a tool of social coercion, a constant reminder of Pip’s status as an unwanted orphan who must show endless gratitude for his mere survival.

Activity 16.2: The Emotional Tug-of-War — Gratitude vs. Regret

Building on Pip’s experiences, I transitioned to Activity 16.2, which explores the broader semantic fields of gratitude and regret across the novel . I ran two comparative concordance searches in CLiC using the multi-term search function: first for words representing gratitude (grateful gracious thankful) and second for words representing regret (remorse remorseful guilt regret) .

This activity beautifully illustrates the power of CLiC's "Subsets" feature. I ran the search under two different parameters:

  1. "Non-quotes" Subset (13 entries, relative frequency of 98.60 pm): This search isolates the retrospective, reflective voice of the adult narrator Pip .

  2. "All text" Subset (14 entries, relative frequency of 75.62 pm): This search includes both narrative reflections and direct character dialogue .

In the "All text" concordance, an extra line appears at Line 10: "...always seen in 'em and always wi' his guilt brought home." Spoken in the heavy dialect of the convict Magwitch, this represents an externalized, legalistic concept of "guilt" . By contrast, the 13 entries in the "Non-quotes" subset reflect Pip’s internal, psychological moral development.


Analyzing the "Non-quotes" concordance reveals that Pip’s psychological journey is defined by a deep, aching self-reproach. As an adult narrator looking back, Pip’s mind is consumed by the "remorse with which my mind dwelt on what my hands had [done]" (Line 1) and "remorseful thoughts" about his "growing change" toward his loyal friend Joe (Line 13) .

When we contrast this with the gratitude search terms, a painful irony emerges. In his childhood, Pip was forced to perform gratitude, yet he felt internally "unthankful" and "ungracious" (Line 4, 6) because he was treated as an outcast . Once he receives his secret fortune, he begins to feel "ashamed of home" (Line 6) . But as an adult, having realized that his wealth came from a convict and that he abandoned the only people who truly loved him, that forced unthankfulness transforms into genuine, self-inflicted "remorse" and "remorseful thoughts" .

By using CLiC to map these semantic fields, we can trace the double-narrator structure of Great Expectations: we see the young Pip experiencing the events in real-time, and the older, remorseful Pip narrating them with moral accountability .

Part 4: The CLiC Activity Book & the Digital Skilling Curriculum

The activities I conducted on Great Expectations are part of a broader pedagogical shift within our postgraduate curriculum at Maharaja Krishnakumarsinhji Bhavnagar University (MKBU) . Under Paper 204: Contemporary Western Theories and Film Studies (specifically Unit 3: Digital Humanities), the "Digital Skilling for Literature Students" initiative aims to equip traditional humanities students with the computational skills required in the information age .

Within this curriculum, the CLiC Activity Book serves as a vital bridge between linguistic science and literary criticism . Historically, language and literature have been taught as separate, often disconnected academic disciplines . The CLiC Activity Book, authored by Michaela Mahlberg, Peter Stockwell, and Viola Wiegand, systematically integrates these two approaches . It illustrates how computer-assisted methods—such as keyword comparison, concordance sorting, and cluster analysis—can directly support close reading and enrich our subjective interpretations of literary texts .

Our instructor hosted an interactive online session to explain the practical mechanics of these activities, guiding us on how to transition from intuitive reading to empirical analysis . During the session, we learned how to navigate the CLiC interface, select target and reference corpora, filter search results using subsets, and utilize the specialized KWICGrouper tool . One of the most valuable practical skills I acquired was learning how to export our findings. After generating a list of keywords or a concordance in CLiC, the interface allows you to save the raw data as a CSV file . By downloading this file and opening it in Excel, we can easily format, sort, filter, and share our data, enabling us to compile empirical evidence and build shared digital readings .

Part 5: Beyond Concordancing — Embracing Voyant Tools and Orange Text Mining

While the CLiC web application is incredibly powerful, it is specifically designed for 19th-century literature . What happens when we want to analyze a modern novel, a collection of digital poetry, or even compare our own essays? During our coursework, we were introduced to two versatile, "limitless canvas" tools: Voyant Tools and Orange . As our instructor noted, their practical applications will be explained in our upcoming online session, but my initial explorations have yielded profound learning outcomes.

Voyant Tools: Instant, Web-Based Visual Hermeneutics

Voyant Tools is a free, web-based text analysis environment designed to facilitate reading and interpretive practices. Unlike CLiC, Voyant is a plug-and-play platform. You simply copy and paste any text or upload a document, and within seconds, Voyant generates a beautiful, interactive dashboard.

My experience using Voyant was highly visual. The platform's view includes several integrated widgets: Cirrus (a dynamic word cloud visualization), Reader (a central panel for traditional close reading), Trends (a line graph tracking the relative frequency of specific words across the document), and Contexts (a concordance view showing search terms in their immediate linguistic environment) .

The key learning outcome from using Voyant is the concept of "distant reading." Voyant makes this accessible. For instance, uploading a novel and tracking the co-occurrence of terms like "death" and "love" across chapters allows us to instantly visualize the emotional and thematic arc of the narrative before we even turn the first page.

Orange: Demystifying Data Mining through Visual Programming

While Voyant provides a fast visual overview, Orange takes us a step further into data science and machine learning. Orange is an open-source data visualization, machine learning, and data mining toolkit. What makes Orange unique is its visual programming interface. Instead of writing complex Python code, users can drag and drop visual "widgets" onto a canvas and link them together to build text processing pipelines.

Using the Orange Text Mining add-on, a typical literary analysis workflow looks like this:


  1. Import Documents widget: Load a corpus of text files.

  2. Preprocess Text widget: Clean the text by performing tokenization, stop-word removal, and lemmatization .

  3. Bag of Words widget: Convert the preprocessed text into a mathematical matrix of frequencies (TF-IDF).

  4. Sentiment Analysis widget: Calculate sentiment scores.

  5. Hierarchical Clustering / Scatter Plot widget: Group documents based on their similarities.

Orange’s widgets demystified these complex computer science concepts. By connecting a "Preprocess Text" widget directly to a "Sentiment Analysis" widget and displaying emotional trajectories on a scatter plot, I realized that data analysis is not a cold, mechanical process. It is a highly structured, visual way of seeing hidden patterns in human language.

By mastering this suite of tools, we can move fluidly between distant reading (using Voyant and Orange to map broad macro-trends across thousands of words) and close reading (using CLiC to examine the micro-details of individual lines) . This dual perspective is the true power of the Digital Humanities.

Conclusion:

The Humanities in the Digital Age — A Personal Reflection

Participating in the "Digital Skilling for Literature Students" coursework has fundamentally transformed my relationship with language and literature . I began this journey with a traditional view of literary criticism—believing that the beauty of a poem or the depth of a novel could only be felt through solitary, silent reading . The idea of subjecting Charles Dickens or William Blake to computer algorithms felt cold, perhaps even sacrilegious .

However, this hands-on experience has shown me that technology is not an enemy of the humanities, but its most powerful collaborator. Whether taking poetry Turing tests , exploring the semantic nuances of "brought up by hand" in CLiC , or building visual preprocessing pipelines in Orange, I have realized that digital tools do not replace the human critic. On the contrary, they demand more of our human capacity for interpretation.

An algorithm can find the 14 instances of "up by hand" in seconds, but it cannot understand the bitter, abusive irony that Mrs. Joe Gargery embeds within those words . An algorithm can calculate that "remorse" is a key term in Pip's retrospective narratorial voice, but it cannot feel the profound moral pain of a grown man realizing he has betrayed the simple blacksmith who loved him . An algorithm can generate a sonnet that mimics Emily Dickinson's style perfectly, but it is the human reader who breathes life into those words, performing the emotional labor of constructing meaning out of random mathematical probabilities .

As Oscar Schwartz beautifully observed, the computer is ultimately a mirror . It does not possess a soul of its own, but it reflects whatever image of humanity we choose to teach it . In an era increasingly dominated by artificial intelligence, our role as students of literature is more critical than ever. We must master these computational tools so that we can actively shape the mirror—ensuring that the reflections of our shared human experience are rich, ethical, diverse, and profoundly meaningful .


No comments:

Post a Comment