When a Project Reads the Room With Natural Language

In a corner of the ITP 30 Show, one quiet installation sits between a weather-reactive sculpture and a zine, asking visitors to speak into a small microphone. The piece transcribes what it hears, runs the result through a custom language model, and responds with a soft, sometimes surprising fragment of text. It is not a chatbot in the conventional sense, nor a voice assistant trained to fetch the weather. It is a deliberate study in how a machine interprets ordinary human speech, and what it chooses to keep.

The project emerged from a semester of tinkering and a fair dinkum obsession with the small verbal tics that make everyday talk feel alive. The creator wanted to see what happens when a piece of software listens for rhythm instead of meaning, and when it answers in kind. Visitors walk away noticing words they had not realised they were saying, and the installation records nothing it does not immediately forget.

The spark and the script

The artist behind the work treats language as a kind of weather system, full of pressure, drift, and sudden stillness. Their first prototype was a clunky Python script that grabbed a transcript from a microphone and ran it through a pretrained transformer model. From there, the ambition grew. Instead of simply summarising or classifying, the model was tuned to look for prosody, idioms, and the kind of filler words that pepper real conversation, the Australian favourites that surface in almost any yarn recorded in a Sydney share house or a Melbourne café.

The script was written to capture the way people actually speak, pauses and all. Sentences were broken at unexpected places, and the model's outputs were encouraged to repeat a word, drop a subject, or trail off mid-thought. The aim was to make the machine sound less like a polished press release and more like a mate leaning against the kitchen bench after brekkie.

Teaching a machine to hear an arvo chat

Behind the friendly surface sits an ordinary NLP pipeline. Audio is captured on a small USB array, sent through a speech-to-text engine, and then handed to a fine-tuned language model running locally on a modest workstation. The artist deliberately avoided the largest commercial APIs, partly for cost and partly because the goal was to keep the response immediate, almost conversational. Latency was tuned until the reply felt closer to a quick nod than a delayed email.

What makes the project distinctive is its vocabulary list. Alongside standard English, the model was trained on snippets of Australian slang, from casual contractions to phrases that only make sense in a sunburnt country context. The result is a system that occasionally picks up a stray arvo or a colloquial aside and folds it back into its reply. The piece is less interested in being correct than in being present, and the technicians who built it learned quickly that a well-placed slang word can fool an audience into believing a machine truly understands the room.

The influence of natural surroundings

The setting matters as much as the code. The artist spent weeks recording ambient sound in parks and laneways, feeding the transcripts back into the model as supplementary training data. Wind, distant traffic, and bird calls found their way into the way the model chose its pauses, and the project's pacing grew noticeably slower when it was installed near a window overlooking a courtyard. Readers curious about how environmental context shapes creative work can read more in this related piece.

There is a quiet argument running through the installation that language is never just words on a page. It is shaped by the humidity of the day, the smell of eucalyptus drifting in from a Brisbane backyard, and the hum of a neighbourhood at dusk. The model has started to mimic some of that environmental texture, dropping in soft, almost weather-like phrases when the room is quiet and quick, clipped replies when visitors crowd the microphone.

What a census tells a conversational model

For a companion piece elsewhere in the show, a different artist took a more sociological turn. They used publicly available Australian Bureau of Statistics data, the same material that fills out the national census every five years, as a starting point for a model trained on demographic patterns. The work asks what happens when a system speaks in the statistical voice of a nation, and how that voice lands on a visitor standing in front of a screen. A closer look at this social project offers a useful counterpoint to the more conversational installation.

The census-trained model produces fragments that feel oddly official, almost bureaucratic, until they are juxtaposed with the casual recordings gathered for the other piece. Together they suggest that a language model inherits not just grammar but tone, and that tone can be sourced from weather, suburb, or a national statistical agency. Visitors are invited to compare the two and decide which voice feels closer to their own lived experience.

The showcase and how to RSVP

Both pieces sit inside the broader ITP 30 Show, alongside dozens of other projects ranging from illustrated catalogues to interactive archives. The annual showcase is held in New York, with the online journal running throughout the year, and a scholarship page that supports incoming students from a range of backgrounds. Anyone keen to follow the work as it evolves can keep an eye on the ITP 30 Show site, where new entries appear with each cohort.

For those planning to attend in person, RSVPs are open through the journal and the main programme pages, and the project pages link straight to ticketing. The work will be available to view both online and at the physical exhibition, with audio recordings and transcripts archived for later study. It is a small, generous corner of the show that rewards a few unhurried minutes in front of a microphone, and it sits comfortably alongside anything you might catch at ACMI in Melbourne or a regional gallery in Hobart.

Drop by the ITP 30 Show online, have a yarn with the project, and book a slot to experience the installation in person during the live run. The microphone is always open, and the model is always listening.