DR McCullochWHISPERING TO THE MACHINE · THREE REGISTERS← WORK
ACADEMICrewrite · reflexive

Introduction

Generative art, where images are produced through algorithms, positions the artist not only as a curator but as an orchestrator, establishing a collaborative bond between the artist and an elusive semi-autonomous system and issuing in a joint artistic endeavour (Pearson, 2011). Within a machine learning (ML) context, algorithms trained on features drawn from vast bodies of existing images can be made to produce pictures in the style of particular artists, styles or movements from algorithmic, written or visual prompts (Mazzone and Elgammal, 2019). To coax these models into producing an image a user supplies text prompts that set in motion a chain of deterministic and probabilistic processes, a practice that has come to be called prompt engineering (PE) (Oppenlaender, 2022).

My practice is one that explores the wider implications of algorithms and ML, and my intention here is to deepen that practice by attending to its very conditions — as Aldouri (2015) puts it, the “art historical legacies, particular institutional power … general institutional forms … and the material conditions of everyday life at a given moment”. I want to be precise from the outset about the position from which I write. I am one located artist: with a particular studio, a particular laptop, a partner who is a Central Saint Martins Fine Art alumnus and, as it turns out, a line of email to a former Turner Prize judge. The knowledge I can produce here is partial, embodied and interested. Rather than treat that partiality as an embarrassment to be disclaimed, I want to treat it as the ground of whatever the paper can honestly claim. What convinces me, what enchants me, what I can afford and who will answer my emails are not noise around the findings; they are the findings.

As an artefact of that practice, I propose that this paper is itself an art object, in that its construction enacts the same processes and methods as the other objects I orchestrate.

To deepen the practice I will learn from the communities that have formed around PE, experiment with popular generative text and image models, reflect on that experiment and then read the encounter back through the academic literature: the place of the artist in relation to PE; the ethical, environmental and social dimensions of ML that are too easily bracketed; and the likely futures of the practice. I organise this around three questions:

  1. Can prompt engineering produce images that have the qualities one would expect from a fine art practice?
  2. What theoretical concerns arise from using ML as part of a critical art practice?
  3. What opportunities will there be for artists with regard to prompt engineering?

Approach
This paper weaves human-machine encounters into its own form. It does so through three ML models; through experiments that respond to the theoretical concerns; and by feeding the output of one model back into my activity and my writing1. I define the approach and methods after the fact, following experimentation rather than before it. The linear format of the paper does not reflect the non-linear way it was made; the structure is a formal construct. Its purpose is threefold: to show the reader the shape a human-machine encounter can take inside a research paper; to develop my own intuition for these models as materials; and to encode, in the very structure of the paper, the problems that large language models pose to research2.

The models are GPT-33, Stable Diffusion4 and DALL-E 25. I chose them because they are available as commercial products with accessible APIs, and because each has drawn significant investment from big tech and finance (Hao, 2020; Krishna, 2022), producing large and active user communities and, less innocently, a dependence on the power structures that constitute techno-capitalism (Parisi, 2019). That choice is worth owning rather than naturalising: to work with these tools at all is to enter an enclosure someone else has already built and to accept its terms. There is no clean outside from which to make this work. I am using instruments I did not design and cannot fully see into, and the honest question is not whether the practice is compromised but what a compromised practice can nonetheless surface.

GPT-3 is an autoregressive language model that produces human-like text (Brown et al., 2020). I use it to generate fragments based on my own writing and on the paper’s references, as a form of automatic writing — a dream state (Bauduin, 2015) outsourced to a machine. This
[1] text will be used to explore how humans and machines can interact and understand each other
[2] response will be used to explore how machines can understand and respond to the text generated by humans6
will appear in a distinct typeface, referenced by a number in brackets, with each new line offering one of the options the algorithm generated; where GPT-3 answers with a number in square brackets I have changed it to round brackets for readability. The intention is to show the reader the texture of the generated text, to develop my own sense of how the model behaves and to stage, within the structure of the paper, the algorithmic ouroboros that Marenko names (Marenko, 2020) — the machine fed on its own and our outputs, doubling back on itself.

DALL-E 2 and Stable Diffusion generate novel images from text and image prompts (Ramesh et al., 2022; Rombach et al., 2021) and are used to respond to the paper’s theoretical concerns and to
[3] provide a visualisation of the potential interactions between humans and machines.
DALL-E 2 was chosen as the newest and most capable model available; Stable Diffusion because it is open-source and therefore more accessible — and, it turns out, more revealing, since an open model lets a practitioner reach the seams and blind-spots a closed one keeps hidden.

I take Merleau-Ponty’s good dialectic as a methodological caution: the arguments I venture are idealisations bound by language, and the gaps between “statements, thesis, antithesis and synthesis” (Merleau-Ponty, Lefort and Lingis, 1992, p.96) are vast and obscured by the very words that bound them. To bridge that gap I work through creative experimentation, as research through art (Earnshaw et al., 2015), treating the models as materials and letting method arise in response to theory, with personal reflection kept as an auto-ethnographic record. Which returns me to the paper’s performance of authority. I will still stage a certain amount of academic drag — the citations, the tables, the measured register — but I want to be clear about what the costume is doing. It is not a disguise meant to pass off shallow understanding as expertise; it is a way of making visible that all knowledge here is dressed, positioned and staged, mine included. The drag foregrounds rather than hides the situated, partial and interested character of what follows. I know these domains rudimentarily and from one angle. Owning that is not false modesty; it is the epistemic condition of the work.

Many theoretical concerns surfaced in an initial review of the literature; I chose a smaller subset (Table 1). This list does not capture the full breadth of what arises from PE but the excluded concerns were very much in mind throughout.

Theoretical Concern Chosen? Reason for choice to include or not
Environmental Impact Mostly a footnote in most ML artist’s practices
Biases and datasets Without understanding this, a practitioner would have a large gap in their practice
Choice of model An integral part of using ML models
From images to video A personal interest with regard to art practice.
Gaming, emergent narrative and interaction A personal interest with regard to art practice.
Art and agency and where the artists sits as a practitioner when using ML Generative work forces us to consider novel ways of thinking about artists and artworks.
Techno-capitalism/determinism Actors within big-tech and finance are large forces within ML model distribution and gain the largest benefit from their use.
Art market and NFTs Theory and research is still quite nascent
Reifying generative work as physical artworks A fascinating concern that would warrant its own in-depth research and exploration
Performance and narrative agency A fascinating concern that would warrant its own in-depth research and exploration
The disappearing artist A bit too theoretically complex and applies to many other forms of art.
Making kin with algorithms A fascinating concern that would warrant its own in-depth research and exploration.

Table 1: The theoretical concerns and reasons for their inclusion in this research paper

One row of that table deserves flagging now, because much of what follows argues against it. Against “Environmental Impact” I wrote that it is “mostly a footnote in most ML artist’s practices”. That is an accurate description of the field and a diagnosis of its failure. I have chosen to refuse the footnote. The material and ecological cost of computation is not a coda to this paper; it is one of its through-lines, and I will keep returning it from the margin to the centre where it belongs.

Prompt Engineering

Algorithms have produced images since the 1960s. The earliest example is AARON, built by the artist Harold Cohen from rules he defined by hand; because those rules were explicit, AARON produced formulaic images with a consistent style and figurative subject (Poltronieri, 2022). With the invention of generative adversarial networks (GANs) in the 2010s and successive innovations in ML, tools emerged that could generate images across many styles and subjects (Fig 1). The architectures differ but the most recent models can build increasingly sophisticated images from prompts made of text, images or both, and can even fill gaps in an image or add elements (Table 2).

Figure 1: Excerpt showing the timeline and output of the current generation of ML models, taken from Cetinic and She’s paper Understanding and Creating Art with AI: Review and Outlook (2021)

Model Name Type of Model Type
Deep Dream7 CNN Text 2 dream, style transfer
Big sleep8 GAN/CLIP Text to image
GANBreeder9 GAN Combining existing images to breed more
DALL-E 210 GPT-3->CLIP Text to image, image & text to image, erase and replace
Midjourney11 CLIP Text to image
Stable Diffusion12 CLIP Text to image, image & text to image, erase and replace
Runway ML13 Multiple Toolset for text to image, image to image, text to 3d texture

Table 2: A Summary of Popular Prompt-based Generative ML Art Tools

To use these tools you supply text, an image or both as a prompt, configure parameters that shape the computation and submit it to be processed into one or more images (Fig 2). The choice of model, parameters and prompt has material, aesthetic and procedural consequences. You can take the output of one model, alter it digitally or in print and feed it back in; by stacking models — GAN-chaining (Plain et al., 2022) — or passing one model’s product to another, you can iterate and transform, following instinct as you would with physical materials, with modifications and visual interventions between each iteration that an algorithm alone would not produce. Treating models as materials, with preferences specific to a practice, seems like
[4] a compelling area of investigation for future research.

Figure 2: Images produced from the prompt “a photograph of an astronaut riding a horse” with different parameters. This is often the default prompt used to compare different SD models and architectures.

These models are new, and the online communities around them are still working out the syntax and semantics of their activity. The accepted term, prompt engineering, came into active use after the release of GPT-3 and first appears around 201914; not everyone is comfortable with it, and some in the design and ML community want a word that better honours the emergent craft of coaxing valued images from a model.

Modifier Impact
Subject terms The desired subject
Style modifiers The style of the image
Image prompts A visual prompt for subject and style
Quality boosters Terms that increase the quality of the output
Repetition Strengthen the association between terms
Magic terms Terms that introduce unpredictable

Table 3: A summary of different modifiers and their impact summarised from A Taxonomy of Prompt Modifiers for Text-To-Image Generation (2022)

Methods here are built collectively. In A Taxonomy of Prompt Modifiers for Text-To-Image Generation (2022) Oppenlaender identifies six modifiers that steer the generated image (Table 3). Communities of tens of thousands share prompts, building up a common repertoire of what produces images they admire (Oppenlaender, 2022). There are search engines for prompts and their images, such as Lexica15, and even a marketplace where prompts are bought and sold16. People interrogate the models, argue over what it means to engineer a prompt17, classify what does and does not work, then share and capitalise on the result.

Reviewing Discord servers18 and Reddit communities19, I would venture that the images these communities produce have a marked homogeneity — the same artists and prompts repeated, converging on the aesthetics of computer games, comics, anime and fantasy art (Fig 3). It is worth naming this homogeneity precisely, because it is not a quirk of taste but a structural tendency. A few high-yield names and modifiers reliably return admired images, so they are replanted endlessly, crowding out the varieties that do not. This is, transposed from soil to latent space, what Vandana Shiva called a monoculture of the mind: a diverse commons of visual culture narrowed to a handful of profitable cultivars, its resilience and its strangeness bred out in favour of predictable yield. And like agricultural monoculture it rests on an enclosure — the uncredited harvesting of countless artists’ work into a seed-bank no one consented to plant. Despite this narrowing, the scale of collective experimentation has produced genuinely sophisticated images, and I would argue that, method aside, they stand visually alongside the work of the artists the model was trained on, with one caveat: on a screen, at low resolution.

Figure 3: Generative images typical of those being produced. Sourced from Lexica.art. The images have no copyright.

Figure 5: Théâtre D'opéra Spatial (Allen, 2022)

But is it Art?
While I was writing this paper an image made with PE (Fig 5) won a prize at the Colorado State Fair, in the digital arts / digitally-manipulated photography category (Roose, 2022); one of the judges, an artist and art teacher, on learning the image was made with ML was untroubled and called it a “beautiful piece” (Metz, 2022). On the strength of that award, could you as a reader answer the first question — whether PE can produce images with the qualities of a fine art practice? I suspect your answer depends heavily on your own position, taste and sense of what a fine art practice is. Personally these images, and those in the PE communities, do not move me past an initial wonder that a machine can make them. So might we hand the decision to someone in a position of cultural authority — an art critic? The suggestion is only half tongue-in-cheek: critics function as de facto gatekeepers of art’s value. But it is worth remembering that a critic is a located authority too, formed inside particular, largely Western and market-linked institutions of taste. To convince one is to learn something about the situation of the critic and the screen, not to settle the nature of the image. Let us ask, then: can PE produce images with the qualities an art critic would expect from a fine art practice?

Can we Convince the Critic? An Experiment. Or, A Paper Within a Paper.
Methodology
The experiment began with a spreadsheet of 25 artists (List A: My Artists), selected from my own photographs taken during gallery visits over the past five years. A Python web crawler gathered text for each artist and Natural Language Processing (NLP) extracted keywords from their biographies — biographies predominantly written by gallerists or curators on gallery or personal webpages. That provenance matters more than it first appears: the keywords that seed the entire pipeline are not raw description but the language of the market’s intermediaries, the terms in which an institution has already agreed to sell each artist. The extraction begins before the model does.

The keywords were then combined by hand with the artists’ names into prompts of the form “list of keywords, by artist”:

sculpture, glass, trash, bottles by Andra Ursuta
installation, sound, politics, perception, information, objects definitions, participation, inter-subjectivity, phenomenology by Shilpa Gupta

This list of my artists was used as a prompt to have GPT-3 complete a further list of 75 artists (List B: GPT-3 Artists) through its sentence-completion.

Stable Diffusion then produced 100 images for each of the 100 prompts drawn from Lists A and B. Each image, initially 512x512 pixels, was upscaled to 1024x1024 with the Real-ESRGAN model.

Commentary, Observations and Insights
The first thing the images revealed was a divergence in appeal. Those from List A had an aesthetic that aligned with my taste; those from List B, mostly well-known artists, did not. This led me to combine prompts from different artists, producing 12 hybrid prompts.

Reviewing the mashups I was mesmerised. Novel relative to any single artist, they reflected my taste in composition and aesthetic; as hybrids they were not identifiable as any one artist, though depending on the combination the image could skew toward one — presumably because that artist carried heavier weightings for the prompt’s words20.

I set the images to save to a cloud drive so I could review them away from the studio, and became addicted to endless scrolling, eventually enchanted by the images from the 12 hybrid prompts. The pleasure was exactly that of the variable-interval reinforcement that underpins gamified social media and gambling machines. It is worth pausing on what that means, because it inverts the usual story of who is training whom. In that moment the machine was not my instrument; my attention was its product. The scroll is a designed capture of the body, tuned to the reward circuitry rather than to judgement, and it works by reducing an open, reflective encounter to compulsion. A critical practice that hopes to resist reduction has to begin by noticing that the reduction is being performed on the practitioner, in the pleasure centres, in real time.

The diffusion model is deterministic: the same prompt and parameters always return the same image, through complex probabilistic computation. My tool saved, alongside each image, a .txt file recording the prompt, parameters and model, so I kept a record of the conditions that produced each image, to reuse and to compare against future architectures.

In total the experiment produced roughly 16,000 images over about 80 hours. My laptop draws 0.125 kW/h; I used around 10 watts of energy costing me about £3.4021 — a trivial sum for 16,000 images (Fig 7). That trivial figure is the most ideologically loaded number in the paper, and it is worth refusing to take at face value. £3.40 is the marginal cost to me. It is cheap precisely because the expensive parts were paid elsewhere and earlier by others: the hundreds of millions spent training the model, the electricity and heat of the data centres, the miners whose landscapes were opened to make the silicon, the labellers who tagged the data for pennies and the artists whose work was harvested without consent to constitute the latent space. £3.40 is not the cost of the images; it is what enclosure feels like from the inside, once the costs have been successfully externalised onto people and places kept out of view. The cheapness is the tell.

Producing 16,000 images this cheaply, this endless fount tickling my hindbrain with dopamine, was curious and unexpected, and I wanted to know whether it was only me. I shared a sample with my partner, a CSM Fine Art alumnus; similarly enchanted, he suggested I compile a selection22 for the art critic and former Turner Prize judge, who replied:

“These are phenomenal! And beautiful - I use that word advisedly because my own efforts with stable diffusion are fairly gruesome: it’s so tempting to just mine it for the uncanny, eerie and post human. But these images are optimistic, aesthetic, imaginative and inspiringly human! Really eye-opening stuff. In fact they make me see the possibilities afresh.

You are right, this is going to transform art.”

To interrogate my own enchantment I printed 100 of the images onto A3 paper and also ran them continuously on a studio screen. I had meant to pin all 100 to the wall but did not — too much work, too little wall, and in any case only 0.6% of the total produced. I let some fall to the floor, where they gathered footprints; a few favourites I pinned up. The printed images lost much of their charm: cheaply laser-printed and, more tellingly, stripped of the texture a screen had seemed to promise. They became flat and less interesting.

Why? If we look at Figure 8, the architecture of the brain is transforming a 2D image into a sculpture made from wood and wool. But the textures, the lighting, the implied spaces and objects are emergent properties of a dataset and an algorithm; because we are so practised at reading 2D representations of 3D things, our own cognitive apparatus generates real texture, physical sculpture and drawn and painted surfaces out of the image. We over-substantiate — we promise, on the evidence of a probabilistic distribution of pixels on a screen, extant places and objects that were never there. Once printed, an image whose promise was a painting, drawing, collage or carving breaks that promise; the flat sheet exposes the gap between the image on the screen and its reality as a digital artefact.

Figure 8: The image produced using the prompt 19. news, history, gossip, oil, painting, lubugo, cloth, narrative, storytelling, bodies, figures, dystopia, sculpture, everyday, assemblage, wood, knitted, natural, craft by Michael Armitage and by Alexandra Bircken and variables Width: 512 Height: 512 Seed: 4728213 Steps: 50 Guidance Scale: 14.0 Prompt Strength: 0.8 Use Face Correction: GFPGANv1.3 Use Upscaling: RealESRGAN_x4plus Sampler: euler_a Negative Prompt: Stable Diffusion Model: C:\stable-diffusion-ui\stable-diffusion\sd-v1-4.ckpt

Reflections on Diversity and Representation
Comparing the two lists (Table 4) exposed a pronounced bias in List B, the list GPT-3 generated, across gender, nationality and race.

Artists’ Characteristic Distribution List A: My Artists Distribution List B: GPT-3 Artists
Gender 56% Male 44% Female 92% Male 8% Female
Nationality Asia: 32% Europe: 28% North America: 28% Africa: 4% South America: 4% 59.09% North America 40.91% Europe 2.27% Asia
Race 33.33% Asian 29.17% White European 20.83% Black 12.5% Latin White: 97.37% Black: 1.32% Asian: 1.32%

Table 4: Distribution of artists by gender, nationality and race. The LLM output had a clear male & eurocentric bias despite the distribution of the artists in the prompt that was used.24

The usual way to read this table is as three separate failures — a gender bias, a race bias and a geographic bias — to be totted up and deplored. I think that reading is too weak, in two specific ways.

First, the categories themselves are not neutral instruments. To sort artists into male or female, into White, Black, Asian or Latin, is already to apply a grid that was drawn under colonial conditions and that does not fit the bodies it is laid over. My own note to the table concedes as much: the numbers are “slightly off” because of dual nationality, mixed heritage and ambiguity, so that some artists fall into two boxes at once24. That slippage is not a rounding error to apologise for; it is the evidence. The bodies exceed the boxes. What the pipeline does is take a messy, plural world of practising artists, force it back through classificatory categories that colonial administration invented and then amplify them — returning the grid to us hardened and cleaned of its exceptions.

Second, the three biases are not additive but compounding and co-constituted. GPT-3 does not merely prefer men, and separately prefer white artists, and separately prefer North America and Europe. It reconstitutes a single figure — the white, Euro-American, male artist — as the default meaning of the word “artist”, against which everyone else registers as a marked deviation. The clearest evidence is in the images themselves: in my experiments the human subjects skewed white European unless I named an artist of another race in the prompt. Whiteness there is not one option among several; it is the unmarked ground that becomes visible only when something is placed against it. My own List A was already diverse — 56% male, and spread across Asian, White European, Black and Latin artists, across five continents. The system did not so much fail to reflect that diversity as actively regress toward the canonical mean, laundering a plural input back into the monoculture. The pipeline is a machine for restoring the centre.

Conclusion
The experiment loosely indicates that one can captivate both the orchestrator and a fine-art audience with the potential of PE — but only on a screen. Printed, the images lose the material qualities we confabulate from a screen on the assumption that we are looking at a photograph of a physical work. On a screen we assume a physical artefact and therefore a maker; in print the illusion shows itself, the promise breaks and both the artist and the art object disappear25.

Where is the Artist?
If we can make images that convince a critic of PE’s potential, where is the boundary a prompt engineer must cross to become an artist? Twentieth-century art history has worried this question about photography, Dadaism, ephemeral work, abstract expressionism, video art, performance art, land art and more; is prompting an algorithm any different? The same charge was once laid against photography — that the photographer merely operated a machine that did the work. Over time the language and practice of photography developed and it became an accepted art, though not everyone who takes photographs is a photographer or an artist: the distinction is made at the interface between a photographer’s self-categorisation as an artist and its confirmation by institutions and audiences. So too, I would argue, with PE: an artist who uses prompt engineering, positions themselves as an artist and has this confirmed by the wider art community is an artist with a particular set of methods, materials and tools. If we take the market as our crude measure, ML works already sell for significant sums in galleries (Christie's, 2018) and in online cryptoart circles (Kent, 2022).

But what if the person prompting the image does not consider themselves an artist, and the work is later sold as art — art without an artist?26 In an early essay the anthropologist Alfred Gell tried to ground a theory of art anthropologically rather than aesthetically, semiotically or art-historically, reconciling Duchamp, Indigenous art and Renaissance painting to ask why nearly all human societies value art objects as they do. We are enchanted, he argued, by the technology of an object’s making — not only the physical craft but the cognitive technology of its ideation, hence Duchamp and the readymade (Gell and Hirsch, 1999). In Art and Agency (1998) Gell goes further: existing theories of art are the dogma and fiction of a cult, and he offers instead a model in which artworks are themselves social agents, existing as a nexus of social relations around a work (van Eck, n.d.). Rather than explain art by its “formal or aesthetic value” or by what artist or gallerist says of it, we should attend to the technologies, cognitive and practical, that enmesh a work in a network of humans, galleries, movements and other agents with their own relationships to the work (Gell, 1998).

That an artwork becomes an agent in its own right, beyond the artist, seems especially apt for prompt-engineered images, which circulate through networks of people who share, iterate and build on both the images and the conditions that produce them. But the nexus cuts the other way too. To make an image with particular qualities you name particular artists in the prompt, and those artists’ work sits inside the dataset the model was trained on; the generated image is squarely within the named artist’s nexus. So has the artist really disappeared? I would argue the artist’s labour is present inside the algorithm and its training data, and that any artist named in a prompt — or whose work helped generate an image — remains an author of it. Here is where I part company with the ethics-footnote framing of this problem. The uncredited absorption of a named artist’s work into a model, harvested without consent and resold as a service, is not a lapse of etiquette to be smoothed over with an acknowledgements page. It is extraction — the enclosure of a commons of human creative labour, the same logic by which a shared inheritance is fenced, renamed as private stock and rented back to the people it was taken from. It is, in the older and more exact sense of the word, a biopiracy applied to culture.

If that is the right description, the remedies follow as questions of who captures the value rather than of good manners. Labour should return a share of the value it creates, or at the very least explicit credit for the image27, and the ability to withdraw from the dataset. Should we encode the artist’s name into these images as metadata, with legal and copyright protection, or design new file formats that use low-energy blockchain technologies to bind artists to the artefacts made from their work? The deeper point is distributive. The value these systems produce was constituted by the very people the systems write out; the platform climbs on a ladder built from their labour and then kicks it away beneath them. Perhaps the prompt engineer who declines the title of artist declines it precisely because they sense this — that the artist has not disappeared but persists, unpaid, thrumming inside the algorithmic black box. I expect we will see technological, curatorial and legal innovation racing to keep pace with the industrial-scale obfuscation on which the most popular models are built.

The Hidden Cost of Machine Learning

In Finding Gaia (2017), Latour reflects on his surprise, as a sociologist, at the relationship between the human and the nonhuman:

“We were still discussing possible links between humans and nonhumans, while in the meantime scientists were inventing … ways to talk about the same thing … on a completely different scale: the “Anthropocene,” the “great acceleration,” “planetary limits,” “geohistory,” “tipping points,” “critical zones,” all these astonishing terms …. terms that scientists had to invent in their attempt to understand this Earth that seems to react to our actions” (Latour, 2017)

That blindspot — that the human and nonhuman are intrinsically intertwined — describes the generative art community too, where most practitioners are only vaguely aware of the physical substrate their work depends on28 while absorbed in its technical, conceptual and aesthetic surface; witness how quickly the community adopted NFTs despite their environmental cost. Algorithms run on silicon, in devices whose physical parts are made through extensive mining that produces environmental devastation in China and across the global south. This geography is not incidental. It is the old colonial map redrawn: the extraction sited where its consequences can be borne by people far from the aesthetic pleasure, the harm exported to the periphery so the centre can enjoy a clean and weightless “cloud”. Onto this is layered a new practice of concealment — “machine-washing”, which “involves misleading information about ethical AI communicated or omitted via words, visuals, or the underlying algorithm of AI itself” (Seele and Schultz, 2022). Nor is it only mining and manufacture: the data centres where computation happens carry a staggering ecological cost — noise pollution, thermal pollution and vast electricity use (Suresh and Guttag, 2021; Monserrate, 2022), with the emissions that follow, and it will only grow as ML becomes ubiquitous. We cannot disregard the
[5] ecological footprint of computation.
[6] impact of algorithms, we must take into account their global reach and responsibility.
[7] fact that algorithms have a very real environmental impact.
and too few artists, in thrall to the surface of ML, are critical of its environmental cost. A critical machine learning practice does not flinch from this dimension; it actively seeks out and explores ways of working that reduce the harm.

It would help those who use these models to understand that they are not the product of some spider-like artificial intelligence crawling across the internet and reading images and text by itself. They require enormous human labour to categorise the data — people working as poorly paid “mechanical turks”, marking features on an unrelenting stream of images (Amoore, 2020; Suresh and Guttag, 2021). The name is more honest than those who use it intend. The original Mechanical Turk was a fraud: a chess-playing automaton with a human being concealed inside, labour disguised as machinery for an audience that preferred the magic. Today’s arrangement is the same trick at planetary scale — hidden human work, disproportionately sited in the global south and disproportionately racialised and gendered, dressed up as autonomous computation. This is time-consuming labour across millions of images, outside the scope of any individual artist (Monserrate, 2022), followed by the vast computational cost of training a model from the data (Incze, 2019), with hundreds of millions spent on the most successful models. Because that labour is beyond any artist, we must rely on models built by tech companies and research institutions — tweaking output with fine-tuning such as textual inversion in Stable Diffusion (Foong, 2022) — and so can only make work that rests on these fraught environmental and ethical conditions. And is it ethical when the tagged images were made by artists included in datasets without their consent, with tech companies paying universities to conduct the training research precisely so as to circumvent intellectual property claims from those whose work is used (Calma, 2021)? The academy is here conscripted to launder the extraction — to give enclosure the alibi of research.

The choice and construction of a dataset shapes a model’s interpretation and output, introducing historical, representation, measurement, learning, evaluation, aggregation and deployment biases before the model ever reaches the public (Fig 9) (Mehrabi et al., 2021). It is worth insisting that these are not one bias sampled at seven points but a bias compounded through a lifecycle, each stage inheriting and amplifying the last, which is exactly why it cannot be patched at the output with a fairness filter. This could be certain popular artists and movements favoured over others, skewing output, or specific races favoured in the production of images — something I met directly, where the subjects skewed white European unless an artist of another race was named. And during my experiments with SD I kept meeting malformed penises whenever a prompt shifted to male nudity, so I made the bias an object of study (Fig 10); on Discord I learned the training data included no genitals. Working directly with an open model, an artist can find the blind-spots and the implicit and explicit biases built into it. There is a biopolitics legible in that comic failure. The model was trained on desire — on scraped pornography among everything else — and then scrubbed of the organs of desire, so that it renders bodies fluent in every respect but the one that was governed out of the data. What a system can and cannot depict is a quiet politics of which bodies are legible, and the grotesque, visible failure stands in for the far more consequential failures that stay invisible — the biases, harder to discover, in models used for prison sentencing, stereotype reinforcement and financial and health discrimination. Artists should not only be aware of these biases; they can capture and communicate them through their practice.

Figure 9: A diagram taken from Suresh and Guttag’s paper showing where bias sits in the machine learning lifecycle (2021)

Fig 10: Blind spots in the dataset: using a pornographic image of a deceased porn star, with a clearly erect penis, and a pornography-related prompt, I prompted a local, uncensored build of Stable Diffusion. The lighting, body-shapes, tattoos and other elements match the source and what pornography leads us to expect; the penis alone is malformed, the gap in the dataset clearly articulated to a viewer.

In Dreams Rewired (Manu Luksch, Martin Reinhart, Thomas Tode), Tilda Swinton narrates a history of communication and technology in which the interests and needs of governments and corporations are shown to be built into how we use it. So when we use ML models we should remember that they are
[8] not developed in a vacuum, but are the result of specific interests and needs, as well as a result of a process that is controlled by a few powerful actors. ML is often presented as a scientific process with the goal of improving efficiency and productivity. However, algorithms are often designed to achieve a specific goal, and these goals are often determined by the interests of a few powerful actors.

Parisi argues that techno-capitalism — the actors who create and mediate our algorithmic world — is generating deterministic futures: algorithms, through their predictive power and their entanglement with our political, social and physical structures, define and reduce future possibilities, so that we must fight back by becoming unpredictable, disruptive and radical (Parisi, 2019). This is the hinge of the whole argument, so let me sharpen it. Determinism is not a property of the mathematics but a political project: it is the promise, sold as inevitability, that the future is already computed and that the only sane response is to optimise inside it. That is the deep grammar of technosolutionism — every problem, including the problems the technology itself creates, awaiting a further technical fix. To resist reduction is to refuse that grammar: to keep open the futures the prediction would close. In a talk, the artist and academic David Benqué (2020) argues for the design practice of diagramming as a way of interrogating algorithms and surfacing their hidden political and social dimensions. An art practice that experiments critically with ML, unearths its hidden costs and works deliberately to confuse algorithmic prediction could contribute to
[9] a collective reimagining of what our futures could look like, that we can hope to subvert and resist the ways in which they are currently used to control and shape our lives
opening new avenues of inquiry and helping to counterbalance the activity of corporate and government actors as we contend for non-dystopian futures.

Reflecting on the concerns we have covered, an artist using generative tools or ML may benefit from asking:

  • Am I complicit in furthering the aims of big tech?
  • Can I train small datasets that I create myself?
  • Can I interrogate the models to capture or translate this uncomfortable truth of the tool I am using?
  • What is the ethical, social and environmental impact of my practice?
  • Who is benefiting from my use of these tools?
  • Can I use these tools as a form of resistance?

The Future

Looking forward, it seems inevitable that these models will advance to generating moving images. There is already research treating moving elements as discrete, trackable units (Fan et al., 2021), and Meta has a text-to-video system29 that produces short clips from prompts; there is also Synthesia, which makes photorealistic avatars that speak a script you have written30. It is easy to imagine tools that produce extended sequences from text, image and video prompts — realising whole scripts, or turning existing footage into scenes with ML-modified mouth movements. Add the models behind deepfakes, which swap faces or change what a subject appears to say, and one can foresee tools that generate video from a script and then tweak body-language, clothing, speech, expression and faces within it; or, conversely, that take clips and a prompt and combine models to transform them into a consistent narrative to iterate on. Suddenly artists have access to video at scales that once demanded enormous money, time and resources. Will we see an explosion in the scale of video art — work that appears to be at the scale of Matthew Barney but is entirely generative? And, conversely, a renewed appreciation for scale, texture and craftsmanship, in the way figurative painting became popular again after a period of decline, abstraction and conceptual exploration (Cullen, 2019), while some creative jobs are lost en masse?

I read that renewed hunger for craft not as nostalgia but as a symptom worth taking seriously: a reaching for the material, the embodied and the situated that the screen cannot deliver, the same over-substantiation that broke in my print experiment surfacing as a collective appetite for the thing itself. It is a counter-current to the monoculture.

Following ML through text, image and video, I propose it will continue into generative gaming and interaction. The theoretical ground is laid — emergent narrative and interactive storytelling (Louchart and Aylett, 2004) — and it will not be long before trained models can generate games or interactive artefacts. But gaming is the largest dataset in the field of interaction, and it sits at what could be argued as the intersection between the military-industrial-complex and the fantasies of adolescent males and
[10] we must be aware that the current AI models are replicating and amplifying existing power structures and behaviours. AI is an extension of our existing structures, not a revolution.
That is precisely the risk to name: the largest interactive corpus encodes a particular structure of desire — militarised, masculinist — and a generative gaming built on it will not transcend that structure but amplify it, unless artists intervene in what is being modelled and for whom. As a medium, gaming has been inherently difficult for artists to work with, as complex animation and moving image are, because the artefacts demand huge teams of specialists and large resources to produce the assets, code, music, lighting, voice acting, scripts and other elements of gameplay. But ML could make this more accessible, letting artists include these technologies in their practice and work at otherwise impossible scales; the question a critical practice keeps asking is whose desires the accessible tools have already been trained to serve.

Conclusion

Given the pace and sophistication of prompt-based ML models and the speed with which PE has been adopted, we should expect continued advances and new methods to emerge, with artists and prompt engineers responding and evolving their practice amid an enormous volume of collaboration and experimentation from entrepreneurs, computer scientists and prompt engineers. Observing the communities and the early research into the taxonomy of prompts, we have seen an increasingly sophisticated practice forming around PE, even if its output is somewhat homogeneous31.

Reflecting on the images made by these communities, the subjectivity of deciding what is and is not fine art is itself a weakness in the first research question. But the experiment inspired and directed work with an ML model that showed images made with PE can excite an art critic, so I would tentatively venture that PE is likely to produce images with the qualities one expects of a fine art practice — with the crucial qualification that this happens on a screen, and that the qualities we grant these images are largely an artefact of our own cognitive apparatus, there being no real texture, scale or other properties that enchant us before physical art objects. Art does not only live on screens. It lives in galleries and on walls and in the hearts of an audience; it has scent and texture and other characteristics that a digitised image lacks. Reading art through Alfred Gell’s anthropology — art without an artist — we find the artist obfuscated but plainly present as an actor inside the training data, and we should expect legal and technological change to follow. But with only the image and none of the rest of the nexus, pure prompt engineers are unlikely to be much valued in a fine art context, a few novelties aside — though they will be valued and rewarded within PE communities, outside the standard fine art modalities32.

Contrary to the pitch from big tech, the work made through ML does not exist in the cloud as digital vapour, a half-remembered cloud of work that lives in the mind of a human artist and inspires them. It is the product of many physical processes with environmental, ethical and social costs, and this is the disconnect between signifier and signified that artists are equipped to expose. My own experiment is a small piece of evidence for exactly this: a set of images that could enchant a former Turner Prize judge on a screen and then, printed and dropped on the floor to gather footprints, give up the illusion entirely. The enchantment and its breaking are the same finding seen twice. Artists with a critical practice can break the enchantment deliberately — can explore and interrogate the material conditions of generative models, make the hidden costs felt, name the extraction as extraction and the enclosure as enclosure, stand with the dispossessed labour thrumming inside the machine and press the distributive question of who is paid and who is written out. Artists are not going to be replaced by machine learning, any more than they were replaced by photography, or photography by film. ML will instead give artists the potential to work at scales and in media that would otherwise be too expensive or difficult, generating many unexpected and surprising outcomes — but in doing so they must refuse the monoculture, contest the marketing machine of techno-capitalism and ensure that the material conditions of algorithmic work are explored, interrogated, challenged and counteracted.

Reference list

Aldouri, H. (2015). On critical art practice: some preliminary reflections. [online] Artblog. Available at: https://www.theartblog.org/2015/11/on-critical-art-practice-some-preliminary-reflections/.

Allen, J.M. (2022). Théâtre D’opéra Spatial. [2022] The New York Times. Available at: https://static01.nyt.com/images/2022/09/01/business/00roose-1/merlin_212276709_3104aef5-3dc4-4288-bb44-9e5624db0b37-superJumbo.jpg?quality=75&auto=webp.

Amoore, L. (2020). Cloud ethics: algorithms and the attributes of ourselves and others. Durham: Duke University Press.

Bauduin, T.M. (2015). The ‘Continuing Misfortune’ of Automatism in Early Surrealism. [online] ScholarWorks@UMass Amherst. Available at: https://scholarworks.umass.edu/cpo/vol4/iss1/10/ [Accessed 03 Oct. 2022].

Benqué, D. (2020). Hybrid Futures: Diagrams for Critical Algorithmic Practice – A talk by David Benqué. [online] www.youtube.com. Available at: https://www.youtube.com/watch?v=c7BFTOzc59U [Accessed 13 Nov. 2022].

Brown, T.B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D.M., Wu, J., Winter, C. and Hesse, C. (2020). Language models are few-shot learners. CoRR, [online] abs/2005.14165. Available at: https://arxiv.org/abs/2005.14165.

Calma, J. (2021). The Climate Controversy Swirling around NFTs. [online] The Verge. Available at: https://www.theverge.com/2021/3/15/22328203/nft-cryptoart-ethereum-blockchain-climate-change.

Cetinic, E. and She, J. (2021). Understanding and Creating Art with AI: Review and Outlook.

Christie's (2018). Is artificial intelligence set to become art’s next medium. [online] Christies.com. Available at: https://www.christies.com/features/A-collaboration-between-two-artists-one-human-one-a-machine-9332-1.aspx.

Cullen, M. (2019). Why has figurative painting become fashionable again? The Spectator. [online] 5 Sep. Available at: https://www.spectator.co.uk/article/why-has-figurative-painting-become-fashionable-again/ [Accessed 13 Nov. 2022].

Earnshaw, R., Liggett, S., Cunningham, S., Heald, K., Thompson, E. and Excell, P. (2015a). Models for research in art, design, and the creative industries. doi:10.1109/ITechA.2015.7317457.

Earnshaw, R., Liggett, S., Cunningham, S., Heald, K., Thompson, E. and Excell, P. (2015b). Models for research in art, design, and the creative industries. doi:10.1109/ITechA.2015.7317457.

Fan, H., Xiong, B., Mangalam, K., Li, Y., Yan, Z., Malik, J. and Feichtenhofer, C. (2021). Multiscale vision transformers. [online] doi:10.48550/ARXIV.2104.11227.

Foong, N.W. (2022). How to Fine-tune Stable Diffusion using Textual Inversion. [online] Medium. Available at: https://towardsdatascience.com/how-to-fine-tune-stable-diffusion-using-textual-inversion-b995d7ecc095 [Accessed 13 Nov. 2022].

Gell, A. (1998). Art and Agency. Clarendon Press.

Gell, A. and Hirsch, E. (1999). The Technology of Enchantment and the Enchantment of Technology. In: The Art of Anthropology: Essays and Diagrams (1st ed.). Routledge.

Hao, K. (2020). The messy, secretive reality behind OpenAI’s bid to save the world. [online] MIT Technology Review. Available at: https://www.technologyreview.com/2020/02/17/844721/ai-openai-moonshot-elon-musk-sam-altman-greg-brockman-messy-secretive-reality/.

Incze, R. (2019). The Cost of Machine Learning Projects. [online] Medium. Available at: https://medium.com/cognifeed/the-cost-of-machine-learning-projects-7ca3aea03a5c.

Kent, C. (2022). NFTs Can Be Artistically Groundbreaking — Meet the Artists and Curators Leading The Way. [online] ARTnews.com. Available at: https://www.artnews.com/list/art-news/artists/what-is-best-nft-art-1234631062/and-the-virtual-joins-the-physical-world/.

Krishna, S. (2022). Stable Diffusion creator Stability AI accelerates open-source AI, raises $101M. [online] VentureBeat. Available at: https://venturebeat.com/ai/stable-diffusion-creator-stability-ai-raises-101m-funding-to-accelerate-open-source-ai/ [Accessed 13 Nov. 2022].

Louchart, S. and Aylett, R. (2004). Narrative theory and emergent interactive narrative. International Journal of Continuing Engineering Education and Lifelong Learning, 14(6), p.506. doi:10.1504/ijceell.2004.006017.

Mazzone, M. and Elgammal, A. (2019a). Art, Creativity, and the Potential of Artificial Intelligence. Arts, 8(1), p.26. doi:10.3390/arts8010026.

Mazzone, M. and Elgammal, A. (2019b). Art, Creativity, and the Potential of Artificial Intelligence. Arts, 8(1). doi:10.3390/arts8010026.

Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K. and Galstyan, A. (2021). A Survey on Bias and Fairness in Machine Learning. ACM Computing Surveys, 54(6), pp.1–35. doi:10.1145/3457607.

Merleau-Ponty, M., Lefort, C. and Lingis, A. (1992). The visible and the invisible. Evanston, Ill. Northwestern University Press, p.96.

Metz, R. (2022). AI won an art contest, and artists are furious. [online] CNN. Available at: https://edition.cnn.com/2022/09/03/tech/ai-art-fair-winner-controversy/index.html.

Monserrate, S.G. (2022). The Staggering Ecological Impacts of Computation and the Cloud. [online] The MIT Press Reader. Available at: https://thereader.mitpress.mit.edu/the-staggering-ecological-impacts-of-computation-and-the-cloud/.

Oppenlaender, J. (2022). A taxonomy of prompt modifiers for text-to-image generation. [online] doi:10.48550/ARXIV.2204.13988.

Parisi, L. (2019). Critical Computation: Digital Automata and General Artificial Thinking. Theory, Culture & Society, 36(2), pp.89–121. doi:10.1177/0263276418818889.

Pearson, M. (2011). Generative art: a practical guide using processing. Shelter Island, Ny: Manning; London.

Plain, A., Eynon, R., Hjorth, I. and Osborne, M.A. (2022). AI and the Arts: How Machine Learning is Changing Artistic Work. Report from the Creative Algorithmic Intelligence Research Project. Oxford Internet Institute, University of Oxford, UK.

Poltronieri, F. (2022). Towards a Symbiotic Future: Art and Creative AI. The Language of Creative AI, pp.29–41. doi:10.1007/978-3-031-10960-7_2.

Ramesh, A., Dhariwal, P., Nichol, A., Chu, C. and Chen, M. (2022). Hierarchical text-conditional image generation with CLIP latents. [online] doi:10.48550/ARXIV.2204.06125.

Rombach, R., Blattmann, A., Lorenz, D., Esser, P. and Björn Ommer (2021). High-resolution image synthesis with latent diffusion models.

Roose, K. (2022). An A.I.-Generated Picture Won an Art Prize. Artists Aren’t Happy. The New York Times. [online] 2 Sep. Available at: https://www.nytimes.com/2022/09/02/technology/ai-artificial-intelligence-artists.html.

Seele, P. and Schultz, M.D. (2022). From Greenwashing to Machinewashing: A Model and Future Directions Derived from Reasoning by Analogy. Journal of Business Ethics, [online] 178(4), pp.1063–1089. doi:10.1007/s10551022050549.

Suresh, H. and Guttag, J. (2021). Understanding Potential Sources of Harm throughout the Machine Learning Life Cycle. MIT Case Studies in Social and Ethical Responsibilities of Computing. doi:10.21428/2c646de5.c16a07bb.

van Eck, C. (n.d.). Gell’s theory of art as agency and living presence response. [online] Leiden University. Available at: https://www.universiteitleiden.nl/en/research/research-projects/humanities/gells-theory-of-art-as-agency-and-living-presence-response [Accessed 1 Nov. 2022].


  1. 1 LARPing as the algorithmic ouroboros.
  2. 2 This won’t follow the approach where I make clear what is and isn’t generative content as the text being input is my own and I see this as no different to using a grammar or spell-checker, or asking advice from a knowledgeable peer and qualifying their view.
  3. 3 https://beta.openai.com/playground
  4. 4 https://github.com/CompVis/stable-diffusion
  5. 5 https://openai.com/dall-e-2/
  6. 6 What is interesting with fragments [1] and [2] is that both impart high-level cognitive behaviours to machines — “understand”, “interact”, “respond” — as if the model were a mind rather than a statistical compression of human labour; the anthropomorphism is not innocent, since it is exactly what lets the human work inside the machine drop out of view. Something to bear in mind as we explore further.
  7. 7 https://deepdreamgenerator.com/
  8. 8 https://github.com/lucidrains/big-sleep
  9. 9 https://github.com/joel-simon/ganbreeder
  10. 10 https://openai.com/dall-e-2/
  11. 11 https://www.midjourney.com/
  12. 12 https://github.com/CompVis/stable-diffusion
  13. 13 https://runwayml.com/
  14. 14 If we review the Wikipedia history for the term (https://en.wikipedia.org/w/index.php?title=Prompt_engineering&action=history) we can see that the term doesn’t appear in Wikipedia until 2021 but in the research literature the first time we see the term used in a paper is 2018 (https://scholar.google.com/scholar?hl=en&as_sdt=0%2C5&q=prompt+engineering&btnG=).
  15. 15 https://lexica.art
  16. 16 https://promptbase.com
  17. 17 https://www.reddit.com/r/midjourney/comments/xqyn9e/prompt_engineering_is_an_art/
  18. 18 https://discord.gg/unstablediffusion; https://discord.gg/ymJMvgJY
  19. 19 https://www.reddit.com/r/StableDiffusion
  20. 20 Prompts are converted into high-dimensional vectors that are used to generate the images so if specific keywords are more weighted to one of the artists the images will skew towards their work that has been encoded in the model. Similarly, if an artist has more work encoded into the model then the weight of the keywords associated with their work will shift the image’s properties. These models are purely probabilistic and my explanation is an over-simplification but roughly allows us to understand the interaction between prompts and artist names.
  21. 21 This calculation was made using https://www.sust-it.net/energy-calculator.php and the highest kw/h value for my laptop’s energy consumption while running the processor and GPU at full performance. It ignores the original cost of training the model and the cumulative cost of its use on a global scale — the very costs the £3.40 is designed to make invisible.
  22. 22 https://drive.google.com/file/d/1w4YZB0tl40g7gw60nbnbT__ADxULIWCL/view?usp=sharing
  23. 24 These categorisations are to illustrate the broad differences in the artists generated by an LLM and my own qualitative impression of this process rather than as a true quantitative representation or study. The numbers are slightly off due to either dual-nationality, ambiguity or mixed-heritage and so artists being included twice in different categories — a slippage that is itself the point, since bodies rarely sit still inside the boxes these categories provide.
  24. 25 See the Wizard of Oz or any of the abrahamic texts.
  25. 26 We already have bread without a baker and candles without a candlestick maker.
  26. 27 I had a chat with a bot based on GPT-3 to explore this idea further and after initial disagreement, it finally agreed that artists used in prompts should receive a share of the profits made from their use.
  27. 28 The vagueness is not the practitioners’ failing so much as the infrastructure’s achievement: the “cloud” is engineered to be imperceptible, its whole rhetorical work being to make computation feel weightless and placeless, so that the substrate can be depended on without ever being felt.
  28. 29 https://makeavideo.studio
  29. 30 https://www.synthesia.io
  30. 31 Although, it isn’t like most national galleries are full of paintings of the wealthy in very similar styles.
  31. 32 Kudos (likes/follows), NFTs, payments for digital artefacts (prompts, models, plugins) or subscription payments to the individual.
ORIGINALthe base · as submitted

Introduction

Generative art, where art is generated using algorithms, positions the artist not just as a curator, but also as an orchestrator. This establishes a collaborative bond between the artist and an elusive autonomous system, resulting in a joint artistic endeavour (Pearson, 2011). Within a machine learning (ML) context, ML algorithms can be utilised to generate images formed from complex computer models trained on features from existing images, producing pictures in the style of specific artists, styles or movements using algorithmic, written or visual prompts (Mazzone and Elgammal, 2019). To use these complex models to produce images, a user will provide text prompts that begin complex deterministic and probabilistic processes through a process called prompt engineering (PE) (Oppenlaender, 2022).

My art practice is one where I explore the wider implications of algorithms and ML and my intention with this paper is to further deepen my practice by considering the very conditions of my work and as Aldouri (2015) succinctly summarises, the “art historical legacies, particular institutional power … general institutional forms … and the material conditions of everyday life at a given moment”. I will also develop approaches and methods to deepen my domain knowledge.

As an artefact generated as part of my practice, I propose that this paper is also an art object, as its construction reflects the same processes and methods as the other art objects that I orchestrate.

To achieve the aim of deepening my practice, I will learn from the communities that have formed around PE and then experiment with popular generative text and prompt-based models and reflect on this. I will then review the academic literature to consider the place of the artist with regard to PE; the broad ethical, environmental, and social dimensions of ML; and anticipate future directions for PE to broadly respond to the following questions:

  1. Can prompt engineering produce images that have the qualities one would expect from a fine art practice?
  2. What theoretical concerns arise from using ML as part of a critical art practice?
  3. What opportunities will there be for artists with regard to prompt engineering?

Approach
This paper will weave human-machine encounters into its format through the use of three ML models; as well as experiments with these models as a response to theoretical concerns; I will also feed output from one model into others that I will inform my activity and writing1; and I will define the approach and methods following experimentation and exploration at the end of this process. The linear format of this paper does not reflect the non-linear approach to its creation, the structure of this paper is a formal construct. The intention of this is to illustrate to the reader the form that an encounter can take within a research paper; develop my own intuition with regard to generative models for my art practice; and finally to encode and exhibit the problems that LLMs pose to research into the very structure of this paper2.

The models I will be using are GPT-33, Stable Diffusion4 and DALL-E 25. They have been chosen because they are available as commercial products with APIs that make them easy to use; as well as significant investment from big-tech and finance (Hao, 2020; Krishna, 2022) resulting in a large and active user community, as well as a dependence on the complex power structures that arguably make up techo-capitalism (Parisi, 2019).

GPT-3 is an autoregressive language model that produces human-like text (Brown et al., 2020) and will be used to generate fragments of text based on my writing and also the references in the paper and as a form of automatic writing that as a dream state (Baudain 2015) I have outsourced to a machine. This
[1] text will be used to explore how humans and machines can interact and understand each other
[2] response will be used to explore how machines can understand and respond to the text generated by humans6
and will be in a slab font and referenced through a number in brackets with each new line providing various options generated by the algorithm; if GPT-3 responds with a number in square brackets I will change this to round brackets for reader comprehension. The intention here is to illustrate to the reader the quality of the text that is generated and also as a practitioner develop my own internal sense of how this model works, as well as simulating the algorithmic ouroboros, coined by Marenko, within the structure of this paper (Marenko, 2020).

DALL-E 2 and Stable Diffusion are models that can generate novel images from text and image prompts (Ramesh et al., 2022; Rombach et al., 2021) and will be used to respond to the theoretical concerns of the paper and
[3] provide a visualisation of the potential interactions between humans and machines. DALL-E 2 was chosen as it is the newest and most advanced model available, while Stable Diffusion was chosen as it is open-source and therefore more accessible. The paper will firstly explore the history of AI and ML, before discussing the concept of the human-machine encounter. It will then go on to describe the three ML algorithms and how they work, before conducting experiments with each model that responds to the theoretical concerns of the paper. Finally, it will reflect on the findings of the paper and what they mean for the future of human-machine encounters.

I also intend my investigation to follow Meleau-Ponty’s definition of the good dialectic where the arguments I document or venture are appreciated as idealisations bound by language and the gaps between “statements, thesis, antithesis and synthesis” (Merleau-Ponty, Lefort and Lingis, 1992, p.96) are vast and obscured by the very words that make up their boundaries. To help bridge this gap, I seek to enrich the words through creative experimentation and as research through art (Earnshaw et al., 2015) where the models I interact with are treated as materials and my methods arise as a response to research and theory, with personal reflections an auto-ethnographic record. The combination of words and the argumentation that I am using should be considered as deceitful, in as much as I am trying to exhibit authority and rigour through academic drag in domains that I have only a rudimentary and shallow understanding.

Although many theoretical concerns were identified as part of an initial review of the academic literature, I chose a smaller subset of these outlined in (Table 1). It is understood that this list in itself does not capture the breadth of theoretical concerns that arise from the practice of PM but were very much in mind while researching this paper.

Theoretical Concern Chosen? Reason for choice to include or not
Environmental Impact Mostly a footnote in most ML artist’s practices
Biases and datasets Without understanding this, a practitioner would have a large gap in their practice
Choice of model An integral part of using ML models
From images to video A personal interest with regard to art practice.
Gaming, emergent narrative and interaction A personal interest with regard to art practice.
Art and agency and where the artists sits as a practitioner when using ML Generative work forces us to consider novel ways of thinking about artists and artworks.
Techno-capitalism/determinism Actors within big-tech and finance are large forces within ML model distribution and gain the largest benefit from their use.
Art market and NFTs Theory and research is still quite nascent
Reifying generative work as physical artworks A fascinating concern that would warrant its own in-depth research and exploration
Performance and narrative agency A fascinating concern that would warrant its own in-depth research and exploration
The disappearing artist A bit too theoretically complex and applies to many other forms of art.
Making kin with algorithms A fascinating concern that would warrant its own in-depth research and exploration.

Table 1: The theoretical concerns and reasons for their inclusion in this research paper

Prompt Engineering

Since the 1960s algorithms have been producing images, the first example of is AARON created by artist Harold Cohen using software he created that uses complex rules defined by Cohen to produce images and with a specific style and figurative subject but due to the explicit nature of the rules, it produced formulaic images with very similar properties (Poltronieri, 2022). In contrast and with the invention of generative adversarial networks (GANs) in the 2010s and further innovations in approaches to ML, we begin to see tools that are able to generate images with different styles and subject matter (Fig 1). The underlying architecture differs but with the most recent models increasingly sophisticated images can be made using prompts made of text, images, or both and can even fill in the gaps in an image or add elements (Table 2).


Figure 1: Excerpt showing the timeline and output of the current generation of ML models taken from Cetinic and She’s paper Understanding and Creating Art with AI: Review and Outlook (2021)

Model Name Type of Model Type
Deep Dream7 CNN Text 2 dream, style transfer
Big sleep8 GAN/CLIP Text to image
GANBreeder9 GAN Combining existing images to breed more
DALL-E 210 GPT-3->CLIP Text to image, image & text to image, erase and replace
Midjourney11 CLIP Text to image
Stable Diffusion12 CLIP Text to image, image & text to image, erase and replace
Runway ML13 Multiple Toolset for text to image, image to image, text to 3d texture

Table 2: A Summary of Popular Prompt-based Generative ML Art Tools

To use the tools, you use text, an image, or both as a prompt, configure specific parameters that will impact the model’s computation and then submit this to be processed to produce one or more images (Fig 2). The choice of model, parameters, and prompt can have material, aesthetic and process implications. You can also choose to take the output of one model, change the image digitally, or print it out and further customise it, with modifications and visual interventions by the artist between each iteration where the choices made by the artist produce an artefact that an algorithm by itself may not. Furthermore, by using multi-layered models and stacking algorithms in a process known as GAN-chaining (Plain et al., 2022) or by taking the product of one model and using it as an input to another, you can iterate, transform and create images, following your instinct as an artist in the same way that you may when experimenting with physical art materials. Within an art practice, the potential of mixing models together and treating them as a material, with personal preferences and approaches that are specific to your practice seems like
[4] a compelling area of investigation for future research.


Figure 2: Images produced from the prompt “a photograph of an astronaut riding a horse” with different parameters. This is often the default prompt used to compare different SD models and architectures.

The recent class of powerful generative models are new and as such, nascent online communities are still figuring out the syntax and semantics of their activity. The currently accepted term is prompt engineering and this became actively used following the release of GPT-3 and was first used as a term around 201914. Not everyone is comfortable with this name and figures within the design and ML community feel that something more poetic that captures the time and burgeoning craft of creating prompts that produce valued images.

Modifier Impact
Subject terms The desired subject
Style modifiers The style of the image
Image prompts A visual prompt for subject and style
Quality boosters Terms that increase the quality of the output
Repetition Strengthen the association between terms
Magic terms Terms that introduce unpredictable

Table 3: A summary of different modifiers and their impact summarised from A Taxonomy of Prompt Modifiers for Text-To-Image Generation (2022)

With prompt engineering, methods are being developed collectively by the communities using these tools and in his paper of A Taxonomy of Prompt Modifiers for Text-To-Image Generation (2022) Oppenlaender identifies six different modifiers that influence the image being generated (Table 3). Online communities of tens of thousands share these prompts, collectively building up methods that can generate images that they collectively admire and value (Oppenlaender, 2022). There are now search engines to explore prompts that have been used and their related images such as Lexica15 and even a marketplace where people buy and sell prompts16. People are interrogating the models, discussing what it means to engineer prompts17, classifying what does and doesn’t work, and then sharing and capitalising on this new market. Reviewing discord servers18 and reddit communities19 I would venture that the images that the communities are producing have a homogeneity to their styles, due to repetition of the same artists and their prompts and as such realise the aesthetics of computer games, comics, anime and fantasy art (Fig 3). Despite this, the scale of experimentation and collaboration has resulted in prompts that produce sophisticated images and I would argue that the images produced, if we ignore their method of production, are visually on an equal footing with those produced by the artists the algorithm is trained on, with one caveat: when viewed on a screen at a low resolution.


Figure 3: Generative images typical of those being produced. Sourced from Lexica.art. The images have no copyright.


Figure 5: Théâtre D'opéra Spatial (Allen, 2022)
But is it Art?
During the composition of this paper an image created using PE (Fig 5) won a local art prize at the Colorado State Fair in their digital arts/digitally-manipulated photography category (Roose, 2022) and one of the judges, an artist and art teacher, on learning that the image was created using ML was happy with his decision and described it as a “beautiful piece” (Metz, 2022). So on the basis of this award and reviewing the image, could you as a reader answer our first question, whether PE can produce images with the qualities expected from a fine art practice? I propose that this answer will be heavily influenced by your own position, taste and perspective regarding the conditions of a fine art practice. Personally, this example and those produced in the communities working with PE do not excite me past an initial wonder that ML can produce such images, so would passing the decision over to someone in a position of cultural authority, an art critic, help us? With a tongue-in-cheek assertion that art critics are the de facto gatekeeper to the value of art. So, let’s ask ourselves “Can PE produce images that have the qualities an art critic would expect from a fine art practice*?”*

Can we Convince the Critic? An Experiment. Or, A Paper Within a Paper.
Methodology
The experiment commenced with the compilation of a spreadsheet containing the names of 25 artists (referred to as List A: My Artists), selected from my personal collection of photographs taken during gallery visits over the last five years. A Python web crawler was used to gather keywords for each artist using Natural Language Processing (NLP) techniques to extract keywords from the artists' biographies, predominantly authored by gallerists or curators on their personal or gallerist webpages.

Subsequently, a manual process was undertaken to create a series of text prompts for GPT-3. These prompts were constructed by combining the extracted keywords with the artists' names, adhering to a specific format that included a list of keywords followed by the artist's name:

sculpture, glass, trash, bottles by Andra Ursuta
installation, sound, politics, perception, information, objects definitions, participation, inter-subjectivity, phenomenology by Shilpa Gupta

This list of my 24 artists was then used as a prompt to generate an additional list of 75 artists (List B: GPT-3 Artists) using GPT-3's sentence completion functionality.

Stable Diffusion was then used to create 100 images for each of the 100 prompts derived from Lists A and B. Each image, initially sized at 512x512 pixels, was upscaled to 1024x1024 pixels using the Real-ERSGAN model.

Commentary, Observations and Insights
The initial analysis of the images generated from Lists A and B revealed a striking divergence in artistic appeal. The images from List A exhibited a compelling aesthetic that aligned with my personal tastes, whereas those from List B, predominantly featuring well-known artists, lacked the same level of intrigue. This observation led to a subsequent process where I combined prompts from different artists, resulting in a collection of 12 unique prompts.

Reviewing the mashup images, I was mesmerised. These images, novel compared to the individual artists themselves, reflected my taste with composition and aesthetic components that excite me and as hybrids of the artists I liked, they weren’t necessarily identifiable as the individual artist, although depending on the artists that were combined some of aesthetic or figurative components of the image were skewed towards one, presumably because one artist had heavier weightings for the words in the prompt20.

I set the images to save to a cloud drive, so I could review the images while away from my studio and became addicted to endlessly scrolling and eventually completely enchanted by the images I was seeing from the 12 prompts created in step 4 of the experiment. I feel that this kind of scrolling and pleasure was similar to the variable-interval schedule reinforcement that underpins gamified interaction with social media content or gambling machines.

It is important to note that the diffusion model will always produce the same image if you feed in the same prompt and parameters, this process is entirely deterministic and based on complex probabilistic algorithms. The tool that I was using outputs both the image and a .txt file with the prompt, parameters and model used in its production and as such I kept a record of the prompts and conditions that generated each image to use in the future and also to compare new iterations of the model and subsequent generative architectures.

In total, the experiment produced roughly 16,000 images over a period of ~80 hours. My laptop runs at 0.125 kw/h so I used about 10 watts of energy costing me ~£3.4021, which is a small sum for 16,000 images (Fig 7).
Producing 16k images for such little cost and effort, this never-ending fount of images that tickled my hindbrain with dopamine, was both curious and unexpected and I wanted to know if it was just me, or if the images were compelling to others. I shared a sample of my favourite images with my partner, a CSM Fine Art alumni, and similarly enchanted he requested that I compile a selection22 to share with the art critic and former Turner Prize judge who responded:

“These are phenomenal! And beautiful - I use that word advisedly because my own efforts with stable diffusion are fairly gruesome: it’s so tempting to just mine it for the uncanny, eerie and post human. But these images are optimistic, aesthetic, imaginative and inspiringly human! Really eye-opening stuff. In fact they make me see the possibilities afresh.

You are right, this is going to transform art.”

To finish this process, and to try and interrogate my enchantment towards the images, I printed out 100 images onto A3 paper and also ran these images continuously on a screen in my studio. I initially wanted to pin 100 images to my studio walls but didn’t as it was too much work and I didn’t have the wall space, this was only 0.6% of the total images produced. I found myself just letting some of the images fall to the floor, where they were covered in foot prints and others that were my favourites I pinned to my wall. The printed images lost a lot of their charm, as they were cheaply laser-printed and also on printing, lost the idea of their extant texture that a screen can promise: they became flat and less interesting.

Thinking about why this happens, if we look at Figure 8, the architecture of our brain is transforming a 2D image into a sculpture made from wood and wool. But the textures, the lighting and the spaces and sculptures are emergent properties of a complex dataset and algorithm, and because we are so used to looking at 2D representations of 3D objects and textures, our own cognitive apparatus is tricking us into generating real texture, physical sculptures and painted, drawn and crafted artworks from the images—we are over-substantiating and promising extant places where the contents and subjects in these images exist based on the probabilistic distribution of pixels on a screen. Once printed out, if the image is one where the promise was a painting, drawing, collage, sculpture etc, the flat printed image breaks this promise and highlights the disparity between the image on the screen and its reality as a digital artefact.

Figure 8: The image produced using the prompt 19. news, history, gossip, oil, painting, lubugo, cloth, narrative, storytelling, bodies, figures, dystopia, sculpture, everyday, assemblage, wood, knitted, natural, craft by Michael Armitage and by Alexandra Bircken and variables Width: 512 Height: 512 Seed: 4728213 Steps: 50 Guidance Scale: 14.0 Prompt Strength: 0.8Use Face Correction: GFPGANv1.3 Use Upscaling: RealESRGAN_x4plusSampler: euler_aNegative Prompt: Stable Diffusion Model: C:\stable-diffusion-ui\stable-diffusion\sd-v1-4.ckpt

Reflections on Diversity and Representation
A comparative analysis of the artist lists (Table 4) revealed a notable bias in the distribution of artists by gender, nationality, and race in List B generated by GPT-3. This discrepancy underscores the limitations of language models in accurately reflecting the diversity of the art world and highlights the need for critical examination of the data sets that inform these models.

Artists’ Characteristic Distribution List A: My Artists Distribution List B: GPT-3 Artists
Gender 56% Male 44% Female 92% Male 8% Female
Nationality Asia: 32% Europe: 28% North America: 28% Africa: 4% South America: 4% 59.09% North America 40.91% Europe 2.27% Asia
Race 33.33% Asian 29.17% White European 20.83% Black 12.5% Latin White: 97.37% Black: 1.32% Asian: 1.32%

Table 4: Distribution of arts by gender, nationality and race. The LLM output had a clear male & eurocentric bias despite the distribution of the artists in the prompt that was used.24

Conclusion
It seems that the experiment loosely indicates that one can captivate both the orchestrator and a fine-art audience to the potential of PE to generate images, but only when observed on a screen. I doubt this would be maintained when these images are printed out, as they lose the material aspects we confabulate from images on a screen through the expectation that the image is a photo of a physical work. You could argue that when we view these images on a screen we assume a physical artefact, and therefore a creator, so that when we print out the image the illusion becomes apparent–the promise is broken and the artist and art object disappear25.

Where is the Artist?
If we’re able to produce images that can convince an art critic of the potential of prompt engineering, where is the boundary that must be crossed for a prompt engineer to transform into an artist then? The 20th century art-historical discourse has been grappling with this question with regard to photography, dadaism, ephemeral work, abstract expressionism, video art, performance art, land art and more, so is prompting an algorithm any different? The same argument was made regarding photography, that the photographer was a user of a tool rather than an artist and that the machine was doing work. But over time, the language and practice of photography developed through experimentation and practice and photography became an accepted form of art, but not all people that take photographs are considered photographers or artists; this distinction is made at the interface between the photographer’s self-categorisation as an artist and its confirmation by those in art institutions and audience. Similarly to photography, I’d argue that if an artist is using prompt engineering within their practice and position themselves as an artist, with this confirmed by the wider art community, then they are an artist with a specific set of methods, materials and tools; and if we just look at the art market as qualification of this, works using ML are already selling for significant sums in galleries (Christie's, 2018) and within online cryptoart circles (Kent, 2022).

But, what if the person prompting a generative artefact does not consider themselves an artist, but it is later sold as an art object—art without an artist?26 In an early paper, the art anthropologist Alfred Gell outlines the foundations of a theory of art from an anthropological perspective (rather than aesthetics, semiotics or art-historical), trying to reconcile Duchamp, Indigenous Art and renaissance painters to try and answer why almost all of human societies appreciate art objects in the way that we do. Gell suggested that we are enchanted by the technology of an object’s creation, not only of the physical craft of its making but also the cognitive technologies involved in its ideation, hence Duchamp with the cognitive technology of the readymade (Gell and Hirsch, 1999). Taking this further, in Art and Agency (1998), Gell argues that existing theories of art are but dogma and fiction created by members of the cult of art and ventures a model of art where artworks in themselves are social agents and exist as a nexus of social relations around a work (van Eck, n.d.). That rather than trying to understand art by its “formal or aesthetic value“, or by what the artist or gallerist says about an artwork, we should try to understand the technologies, cognitive and practical, that enmesh individual artworks in a network of humans, galleries, movements and other agents with their own relationships with the work itself (Gell, 1998). The idea that an artwork itself once produced becomes an agent, outside of the artist, is something that I feel becomes particularly relevant in images generated through prompt engineering—where the image itself is part of networks of relationships with people who then share, iterate and build upon these images and the prompts and algorithmic conditions that precipitate them. And furthermore, to produce generative images of with specific qualities, specific artists are named in the prompts, as well as artists work included in the dataset that makes up the model, so the images generated seem well within the named artist’s nexus and as such, has the artist really disappeared? I would argue that the labour of the artist exists within the algorithm and the dataset from which it was trained, and that any artist named within a prompt (or whose work has contributed to the generation of an image) remains the author of any image that is produced. Labour should result in a share of the financial gain made from their work or at a minimum explicit credit for the image27 and the opportunity to remove themselves from the dataset? Should we encode the artist’s name into these images as meta-data, with legal and copyright protections, or create new file formats that use low-energy blockchain technologies to connect artists with the artefacts produced from their work? Our original example of the prompt engineer who produces a generative image using a model trained on the work of others and who doesn’t claim to be an artist, maybe it is because they understand that the artist hasn't disappeared, the artist exists as unpaid labour that thrums inside the algorithmic black box. I imagine we will see new technological, curatorial and legal innovations to keep pace with the industrial-scale obfuscation that is a foundation of many of the most popular ML models.

The Hidden Cost of Machine Learning

In Finding Gaia (2017), Latour reflects on his surprise as a sociologist with regard to the relationship between the human and nonhuman:

“We were still discussing possible links between humans and nonhumans, while in the meantime scientists were inventing … ways to talk about the same thing … on a completely different scale: the “Anthropocene,” the “great acceleration,” “planetary limits,” 9“geohistory,” “tipping points,” “critical zones,” all these astonishing terms …. terms that scientists had to invent in their attempt to understand this Earth that seems to react to our actions” (Latour, 2017)

This blindspot, that the human and nonhuman are intrinsically intertwined, could also be applied to the generative art community where most practitioners are vaguely aware of the physical substrate their work relies on28, but are otherwise preoccupied with their exploration of the technical, conceptual and aesthetic components of ML: we just have to look at the quick adoption of NFTs by this community despite its environmental costs. Algorithms run on silicon and in devices made up of physical parts that are manufactured through extensive mining activity, producing environmental devastation in China and the global south; where we have companies intentionally orchestrating greenwashing and a new practice, known as “machine-washing, that “involves misleading information about ethical AI communicated or omitted via words, visuals, or the underlying algorithm of AI itself” (Seele and Schultz, 2022). It is not only mining and production of the substrates of computation that has an impact, the data centres where computation occurs also have a staggering ecological impact, including noise pollution, thermal pollution and vast use of electricity (Suresh and Guttag, 2021) with its corresponding emissions, and this will only increase as ML becomes increasingly ubiquitous. As such, we cannot disregard the
[5] ecological footprint of computation.
[6] impact of algorithms, we must take into account their global reach and responsibility.
[7] fact that algorithms have a very real environmental impact.
and far too few artists, in thrall to the surface of ML, are critical of its environmental cost. I would argue that a critical machine learning practice does not shy away from this dimension and instead actively seeks out and explores alternative ways of working that reduce its impact.

Those using ML models might benefit from the understanding that the models they use are not the product of some spider-like artificial intelligence crawling across the internet and processing images and text; they require enormous amounts of human labour to categorise datasets with the people acting as poorly paid “mechanical turks” where they are marking features on a constraint stream of images (Amoore, 2020; Suresh and Guttag, 2021). This is an incredibly time-consuming process and involves working with millions of images, which is outside the scope of most artists (Monserrate, 2022) with massive costs with regard to computation to then process and generate a model from these datasets (Incze, 2019), with £100s of millions spent on the most popular and successful models. This manual labour is well outside the scope of any individual artist, so an artist must rely on existing models created by tech companies and research institutions, with the possibility of tweaking output using fine-tuning techniques such as textual inversion in the case of stable diffusion (Foong, 2022), so they can only produce work that relies on fraught environmental and ethical conditions. And is it ethical when the images that are tagged have been created through the labour of artists that have been included in datasets without their consent, with tech companies paying universities to conduct the research to train the models to intentionally circumvent intellectual property claims from those whose work is used to train their models (Calma, 2021)?

The choice of dataset and the way that a dataset is compiled to train an ML model has an impact on its interpretation and output and these introduce historical, representation, measurement, learning, evaluation, aggregation and deployment biases before the model hits the world and is used by the general public (Fig 9) (Mehrabi et al., 2021). This could be certain popular artists and movements being favoured over others in datasets, skewing their output or specific races being favoured over others in the production of images, something that I experienced in my experiments where the subjects skewed white European unless an artist of a different race was used as a prompt. Furthermore, during my experimentation with SD I came across malformed penises whenever the prompt shifted to one where an image had male nudity, so experimented with this bias (Fig 10) and on further investigation on Discord learned that the datasets used to train the model did not include genitals. So working directly with a model as an artist you can discover blind-spots and the implicit and explicit biases contained within the models themselves. Potentially, you could produce work that communicates this—the grotesque penises seem to capture how horrific other blind spots can be in ML, biases that aren’t so easy to discover in ML models that are used for prison sentencing, stereotype reinforcement, financial and health discrimination. Artists should not only be aware of these biases, but can also capture and communicate them through their practice.


Figure 9: A diagram taken from Suresh and Guttag’s paper showing where bias sits in the machine learning lifecycle (2021)


Fig 10: Blind spots in the dataset: in this experiment I used a pornographic image of a deceased porn star, with a clear erect penis and a prompt related to pornography and prompted a local version of Stable Diffusion that does not censor nudity. The lighting, body-shapes, tattoos and other elements that make up the image are inline with the source image and what would be expected in pornography but the penis is malformed, the gap in the dataset clearly articulated to a viewer.

In Dreams Rewired (Manu Luksch, Martin Reinhart, Thomas Tode), Tilda Swinton narrates a history of communication and technology that captures how governments and corporations interests and needs are implicit aspects of how we use and interact with technology. So if we are using ML models, we need to remember that they are
[8] not developed in a vacuum, but are the result of specific interests and needs, as well as a result of a process that is controlled by a few powerful actors. ML is often presented as a scientific process with the goal of improving efficiency and productivity. However, algorithms are often designed to achieve a specific goal, and these goals are often determined by the interests of a few powerful actors.

Parisi argues that techno-capitalism, made up of actors that create and mediate our technological (algorithmic) world, are generating deterministic futures. The determinism comes from algorithms that through their predictive power and entanglement with our political, social and physical structures define and reduce future possibilities, and as such we need to fight against this by becoming unpredictable, disruptive and radical (Parisi, 2019). In a talk by artist and academic David Benqué (2020), he argues that we should use the design practice of diagramming as a way of interrogating algorithms and surfacing hidden political and social dimensions. So similarly, an art practice that critically experiments with ML and unearths its hidden costs and is adopted to intentionally confuse algorithmic prediction could contribute to
[9] a collective reimagining of what our futures could look like, that we can hope to subvert and resist the ways in which they are currently used to control and shape our lives
opening up new avenues of inquiry and working to balance out the activity of corporate and government actors as we wrestle for non-dystopian futures.

Reflecting on the concerns we’ve covered, an artist who is using generative tools or ML may benefit from asking themselves:

  • Am I complicit in furthering the aims of big-tech?
  • Can I train small datasets that I create myself?
  • Can I interrogate the models to capture or translate this uncomfortable truth of the tool that I am using?
  • What is the ethical, social and environmental impact of my practice
  • Who is benefiting from my use of these tools?
  • Can I use these tools as a form of resistance?

The Future

Looking forward, it seems inevitable that ML models will advance to generating moving images. There is already research outlining approaches that track, generate and treat moving elements as individual discrete units (Fan et al., 2021) with Meta ready to launch a new text-to-video system29 that can produce short clips from text prompts. There is also an existing tool called Synthesia that can create photo-realistic avatars that speak and follow a script you have written30. Looking at the output of Meta’s tool and that of Synthesia, it seems very likely that we will have tools that are able to produce extended video sequence from text, image and video prompts, potentially realising entire scripts; or turning existing footage into scenes and script with mouth movements modified using ML. If we consider other ML models that work with video, such as those used in deepfakes where a person’s expression is changed, or their face is swapped, or where footage can be modified so that the subject speaks different words, it becomes easy to imagine tools where you generate video from a script and then tweak body-language, clothing, speech, expression, faces and other discrete items within the video. Or conversely, a process where you feed a tool video clips and a prompt, where multiple models combine and transform your input into a consistent narrative that you then tweak and then use as the source for further experimentation. Suddenly, we see artist’s with tools that allow them to work with video that would otherwise take a vast amount of money, time and resources to produce. Will we see an explosion in the scale of video art, with artists producing work that appears to be at the scale of Matthew Barney but is entirely generative? And conversely, will we also see a greater appreciation for scale, texture and craftsmanship in the same way that figurative art and painting have become popular after a period of decline, abstraction and conceptual exploration (Cullen, 2019) and as some jobs in creative industries are lost en masse?

Reflecting on ML as it follows a trajectory through different media, from text, image and video, I propose that we should expect it to continue into generative gaming and interaction. With the theoretical foundations in place for interaction through the gaming world and concepts such as emergent narrative and interactive storytelling (Louchart and Aylett, 2004), it shouldn’t be too long before we see complex trained models that are able to produce games or interactive artefacts. But, gaming makes up the largest dataset within the field of interaction and exists in what could be argued as the intersection between the military-industrial-complex and the fantasies of adolescent males and
[10] we must be aware that the current AI models are replicating and amplifying existing power structures and behaviours. AI is an extension of our existing structures, not a revolution.
As artists, gaming technologies as a medium are inherently difficult to work with in a similar way that complex animation or moving image is, in that the artefacts produced involve the collaboration of huge teams of individual experts and large amounts of resources to produce the assets, code, music, lighting, voice acting, scripts and other elements of gameplay and interaction. But, with machine learning, this process could become more accessible as ML tools could be used to replace this activity, supporting the opportunity for artists to include these technologies as part of their practice and enabling work at scales that would be impossible otherwise.

Conclusion

Through the pace of change and sophistication we observe with prompt-based machine learning models and the speed at which prompt engineering has been adopted, we should expect continued advances and new methods to rapidly emerge with artists and prompt engineer’s responding and evolving their practice due to the sheer scale of collaboration and experimentation from entrepreneurs, computer scientists and prompt engineers. Observing the online communities that have formed around these models—as well as early research into the taxonomy of prompts—we have seen that there is an increasingly sophisticated practice forming around PE, even if its output is somewhat homogenous31. Reflecting on the images created by these communities, it is clear that the subjectivity that comes with deciding what is and isn’t fine art reveals a weakness in the first research question. But, the investigation inspired and directed experimentation with an ML model that revealed that images produced using PE can excite an art critic and as such could tentatively venture that prompt engineering is likely to produce images that have the qualities one would expect from a fine art practice. However, these images are on a screen and the qualities we impart on these images are mostly an artefact of our own cognitive apparatus as there is no texture, scale and other properties that enchant us when we see physical art objects. Art doesn’t just exist on screens, it exists in galleries and on walls and in the hearts of audience—it has scent, texture and other characteristics that are missing from a digitised image—and using Alfred Gell’s art anthropology as a way of thinking about art without artists, we see that although the artist has been obfuscated, they are still clearly present as an actor within the dataset that is used to train a generative model that and it becomes clear that we are likely to see legal and technological changes that will impact the models that underpin PE but also, with just the image and none of the other parts of the nexus it is unlikely that pure prompt engineers are going to be particularly valued in a fine art context, outside a few novelties; but who will be valued within PE online communities and rewarded outside of standard fine art modalities32. Contrary to the pitch from big-tech, the work produced through ML does not exist in the cloud as digital vapour, a cloud of half-remembered work that exists within the mind of a human artist and inspires their work, it is the product of many physical processes that have environmental, ethical and social costs⁠—and as artists we can challenge this disconnect between the signifier and signified. Artists with a critical practice have the skills to explore and interrogate the material conditions of generative models⁠—or even manifest this as a radical art practice that directly disrupts, co-opts and challenges the position of those who intentionally obfuscate its real-world impact. Artists are not going to be replaced by machine learning, as they were not replaced by photography and photography was not replaced by film. Instead, ML will provide the potential for artists to work at scales and in mediums that would be otherwise too expensive or difficult to work with, generating many unexpected and surprising outcomes, but in doing so they must challenge the marketing machine of techno-capitalism and ensure that the material conditions of algorithmic work are explored, interrogated, challenged and counteracted.

Reference list

Aldouri, H. (2015). On critical art practice: some preliminary reflections. [online] Artblog. Available at: https://www.theartblog.org/2015/11/on-critical-art-practice-some-preliminary-reflections/.

Allen, J.M. (2022). Théâtre D’opéra Spatial. [2022] The New York Times. Available at: https://static01.nyt.com/images/2022/09/01/business/00roose-1/merlin\_212276709\_3104aef5-3dc4-4288-bb44-9e5624db0b37-superJumbo.jpg?quality=75\&auto=webp.

Amoore, L. (2020). Cloud ethics : algorithms and the attributes of ourselves and others. Durham: Duke University Press.

Bauduin, T.M. (2015). The ‘Continuing Misfortune’ of Automatism in Early Surrealism. [online] ScholarWorks@UMass Amherst. Available at: https://scholarworks.umass.edu/cpo/vol4/iss1/10/ [Accessed 03 Oct. 2022].‌

Benqué, D. (2020). Hybrid Futures: Diagrams for Critical Algorithmic Practice – A talk by David Benqué. [online] www.youtube.com. Available at: https://www.youtube.com/watch?v=c7BFTOzc59U [Accessed 13 Nov. 2022].

Brown, T.B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D.M., Wu, J., Winter, C. and Hesse, C. (2020). Language models are few-shot learners. CoRR, [online] abs/2005.14165. Available at: https://arxiv.org/abs/2005.14165.

Calma, J. (2021). The Climate Controversy Swirling around NFTs. [online] The Verge. Available at: https://www.theverge.com/2021/3/15/22328203/nft-cryptoart-ethereum-blockchain-climate-change.

Cetinic, E. and She, J. (2021). Understanding and Creating Art with AI: Review and Outlook.

Christie's (2018). Is artificial intelligence set to become art’s next medium. [online] Christies.com. Available at: https://www.christies.com/features/A-collaboration-between-two-artists-one-human-one-a-machine-9332-1.aspx.

Cullen, M. (2019). Why has figurative painting become fashionable again? The Spectator. [online] 5 Sep. Available at: https://www.spectator.co.uk/article/why-has-figurative-painting-become-fashionable-again/ [Accessed 13 Nov. 2022].

Earnshaw, R., Liggett, S., Cunningham, S., Heald, K., Thompson, E. and Excell, P. (2015a). Models for research in art, design, and the creative industries. doi:10.1109/ITechA.2015.7317457.

Earnshaw, R., Liggett, S., Cunningham, S., Heald, K., Thompson, E. and Excell, P. (2015b). Models for research in art, design, and the creative industries. doi:10.1109/ITechA.2015.7317457.

Fan, H., Xiong, B., Mangalam, K., Li, Y., Yan, Z., Malik, J. and Feichtenhofer, C. (2021). Multiscale vision transformers. [online] doi:10.48550/ARXIV.2104.11227.

Foong, N.W. (2022). How to Fine-tune Stable Diffusion using Textual Inversion. [online] Medium. Available at: https://towardsdatascience.com/how-to-fine-tune-stable-diffusion-using-textual-inversion-b995d7ecc095 [Accessed 13 Nov. 2022].

Gell, A. (1998). Art and Agency. Clarendon Press.

Gell, A. and Hirsch, E. (1999). The Technology of Enchantment and the Enchantment of Technology. In: The Art of Anthropology: Essays and Diagrams (1st ed.). Routledge.

Hao, K. (2020). The messy, secretive reality behind OpenAI’s bid to save the world. [online] MIT Technology Review. Available at: https://www.technologyreview.com/2020/02/17/844721/ai-openai-moonshot-elon-musk-sam-altman-greg-brockman-messy-secretive-reality/.

Incze, R. (2019). The Cost of Machine Learning Projects. [online] Medium. Available at: https://medium.com/cognifeed/the-cost-of-machine-learning-projects-7ca3aea03a5c.

Kent, C. (2022). NFTs Can Be Artistically Groundbreaking — Meet the Artists and Curators Leading The Way. [online] ARTnews.com. Available at: https://www.artnews.com/list/art-news/artists/what-is-best-nft-art-1234631062/and-the-virtual-joins-the-physical-world/.

Krishna, S. (2022). Stable Diffusion creator Stability AI accelerates open-source AI, raises $101M. [online] VentureBeat. Available at: https://venturebeat.com/ai/stable-diffusion-creator-stability-ai-raises-101m-funding-to-accelerate-open-source-ai/ [Accessed 13 Nov. 2022].

Louchart, S. and Aylett, R. (2004). Narrative theory and emergent interactive narrative. International Journal of Continuing Engineering Education and Lifelong Learning, 14(6), p.506. doi:10.1504/ijceell.2004.006017.

Mazzone, M. and Elgammal, A. (2019a). Art, Creativity, and the Potential of Artificial Intelligence. Arts, 8(1), p.26. doi:10.3390/arts8010026.

Mazzone, M. and Elgammal, A. (2019b). Art, Creativity, and the Potential of Artificial Intelligence. Arts, 8(1). doi:10.3390/arts8010026.

Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K. and Galstyan, A. (2021). A Survey on Bias and Fairness in Machine Learning. ACM Computing Surveys, 54(6), pp.1–35. doi:10.1145/3457607.

Merleau-Ponty, M., Lefort, C. and Lingis, A. (1992). The visible and the invisible. Evanston, Ill. Northwestern University Press, p.96.

Metz, R. (2022). AI won an art contest, and artists are furious. [online] CNN. Available at: https://edition.cnn.com/2022/09/03/tech/ai-art-fair-winner-controversy/index.html.

Monserrate, S.G. (2022). The Staggering Ecological Impacts of Computation and the Cloud. [online] The MIT Press Reader. Available at: https://thereader.mitpress.mit.edu/the-staggering-ecological-impacts-of-computation-and-the-cloud/.

Oppenlaender, J. (2022). A taxonomy of prompt modifiers for text-to-image generation. [online] doi:10.48550/ARXIV.2204.13988.

Parisi, L. (2019). Critical Computation: Digital Automata and General Artificial Thinking. Theory, Culture & Society, 36(2), pp.89–121. doi:10.1177/0263276418818889.

Pearson, M. (2011). Generative art : a practical guide using processing. Shelter Island, Ny: Manning ; London.

Plain, A., Eynon, R., Hjorth, I. and Osborne, M.A. (2022). AI and the Arts: How Machine Learning is Changing Artistic Work. Report from the Creative Algorithmic Intelligence Research Project. Oxford Internet Institute, University of Oxford, UK.

Poltronieri, F. (2022). Towards a Symbiotic Future: Art and Creative AI. The Language of Creative AI, pp.29–41. doi:10.1007/978-3-031-10960-7_2.

Ramesh, A., Dhariwal, P., Nichol, A., Chu, C. and Chen, M. (2022). Hierarchical text-conditional image generation with CLIP latents. [online] doi:10.48550/ARXIV.2204.06125.

Rombach, R., Blattmann, A., Lorenz, D., Esser, P. and Björn Ommer (2021). High-resolution image synthesis with latent diffusion models.

Roose, K. (2022). An A.I.-Generated Picture Won an Art Prize. Artists Aren’t Happy. The New York Times. [online] 2 Sep. Available at: https://www.nytimes.com/2022/09/02/technology/ai-artificial-intelligence-artists.html.

Seele, P. and Schultz, M.D. (2022). From Greenwashing to Machinewashing: A Model and Future Directions Derived from Reasoning by Analogy. Journal of Business Ethics, [online] 178(4), pp.1063–1089. doi:10.1007/s10551022050549.

Suresh, H. and Guttag, J. (2021). Understanding Potential Sources of Harm throughout the Machine Learning Life Cycle. MIT Case Studies in Social and Ethical Responsibilities of Computing. doi:10.21428/2c646de5.c16a07bb.

van Eck, C. (n.d.). Gell’s theory of art as agency and living presence response. [online] Leiden University. Available at: https://www.universiteitleiden.nl/en/research/research-projects/humanities/gells-theory-of-art-as-agency-and-living-presence-response [Accessed 1 Nov. 2022].


  1. 1 LARPing as the algorithmic ouroboros.
  2. 2 This won’t follow the approach where I make clear what is and isn’t generative content as the text being input is my own and I see this as no different to using a grammar or spell-checker, or asking advice from a knowledgeable peer and qualifying their view
  3. 3 https://beta.openai.com/playground
  4. 4 https://github.com/CompVis/stable-diffusion
  5. 5 https://openai.com/dall-e-2/
  6. 6 What is interesting with fragments [1] and [2] is that both fragments impart high-level cognitive behaviours to machines by using adjectives such as “understand”, “interact” and “respond”, something to bear in mind as we explore further
  7. 7 https://deepdreamgenerator.com/
  8. 8 https://github.com/lucidrains/big-sleep
  9. 9 https://github.com/joel-simon/ganbreeder
  10. 10 https://openai.com/dall-e-2/
  11. 11 https://www.midjourney.com/
  12. 12 https://github.com/CompVis/stable-diffusion
  13. 13 https://runwayml.com/
  14. 14 If we review the wikipedia history for the term at Wikipedia (https://en.wikipedia.org/w/index.php?title=Prompt_engineering&action=history) we can see that the term doesn’t appear in Wikipedia until 2021 but in the research literature the first time we see the term used in a paper is 2018 (https://scholar.google.com/scholar?hl=en\&as\_sdt=0%2C5\&q=prompt+engineering\&btnG=)
  15. 15 https://lexica.art
  16. 16 https://promptbase.com
  17. 17 https://www.reddit.com/r/midjourney/comments/xqyn9e/prompt\_engineering\_is\_an\_art/
  18. 18 https://discord.gg/unstablediffusion; https://discord.gg/ymJMvgJY,
  19. 19 https://www.reddit.com/r/StableDiffusion;
  20. 20 Prompts are converted into high-dimensional vectors that are used to generate the images so if specific keywords are more weighted to one of the artists the images will skew towards their work that has been encoded in the model. Similarly, if an artist has more work encoded into the model then the weight of the keywords associated with their work will shift the image’s properties. These models are purely probabilistic and my explanation is an over-simplification but roughly allows us to understand the interaction between prompts and artist names.
  21. 21 This calculation was made using https://www.sust-it.net/energy-calculator.php and the highest kw/h value for my laptop’s energy consumption while running the processor and GPU at full performance. It ignores the original cost of training the model and the cumulative cost of its use on a global scale.
  22. 22 https://drive.google.com/file/d/1w4YZB0tl40g7gw60nbnbT\_\_ADxULIWCL/view?usp=sharing
  23. 24 These categorisations are to illustrate the broad differences in the artists generated by an LLM and my own qualitative impression of this process rather than as a true quantitative representation or study. The numbers are slightly off due to either dual-nationality, ambiguity or mixed-heritage and so artists being included twice in different categories.
  24. 25 See the Wizard of Oz or any of the abrahamic texts.
  25. 26 We already have bread without a baker and candles without a candlestick maker.
  26. 27 I had a chat with a bot based on GPT-3 to explore this idea further and after initial disagreement, it finally agreed that artists used in prompts should receive a share of the profits made from their use.
  27. 28 Reference to the paper on what a generative artist should consider
  28. 29 https://makeavideo.studio
  29. 30 https://www.synthesia.io
  30. 31 Although, it isn’t like most national galleries are full of paintings of the wealthy in very similar styles.
  31. 32 Kudos (likes/follows), NFTs, payments for digital artefacts (prompts, models, plugins) or subscription payments to the individual.
LITERARYreimagining · essay

An account of eighty hours, sixteen thousand images and three pounds forty

The Machine Doesn't Stop

For eighty hours the laptop did not stop. It sat on the desk with its fan up like a small aggrieved animal, ten watts of it, warm to the back of the hand, and every few seconds it laid another image into the folder — the way a hen that has lost its mind will keep laying, indifferent and productive and slightly obscene. I had told it what to want. That was the whole of my contribution. I whispered to it, and it went off into its probabilities, and it came back with a picture, and then another picture, and then fifteen thousand nine hundred and ninety-eight more.

I want to be honest at the outset that this account is itself the work. Not a report of the thing but the thing — an artefact made the same way I make the others, by orchestrating a system I don't fully control and taking the credit. The machine and I wrote it together. Where you see it speak — in a slab-serif italic, with a number in a bracket — that is the ouroboros, the snake I fed its own tail: the model finishing sentences I had started, so that the essay quietly eats and generates itself as you read.1 I'm not going to flag every seam. What I put in was mine, and letting a machine complete a line feels no worse than a spell-checker, or asking a clever friend and keeping the half you like.2 The stitches are showing. That is on purpose.

Here is how it speaks. I write, this text will be used to explore how humans and machines can — and the machine takes the sentence out of my hands: [1] interact and understand each other. I offer it another. This response will be used to explore how machines can[2] understand and respond to the text generated by humans. Note the verbs it reaches for, unbidden. Understand. Interact. Respond. It has decided, on my behalf and without being asked, that it has an interior.3 It does not. But it will describe one, fluently, all night, at a fraction of a penny.

There were three of them on the bench. GPT-3, which produces text so plausibly human that its own makers slip and call it understanding. Stable Diffusion, open-source and therefore mine to run at home until the small hours, which turns a line of words into a picture and keeps the receipt. And DALL-E 2, the newest and the cleverest and the most locked-down of the three, kept behind its makers' glass where I could look but not tinker. Between them, I told myself, in the grant-application voice one slips into, they would [3] provide a visualisation of the potential interactions between humans and machines — a sentence the machine finished for me without a flicker, having heard it, or something near it, ten thousand times before. Big money stands behind all of them, and behind the money the whole apparatus that a certain kind of theorist means when they say techno-capitalism, though we'll come to that, in the mine and in the shed and in the noise.

I went in carrying three questions, the way you take three coins to a well. Can a prompt — a whispered instruction to a machine — produce an image with the qualities you'd want from a fine-art practice? What does it cost you, theoretically and otherwise, to make one of these models part of a critical practice? And what, if anything, is actually in it for artists? I'll tell you now that I came out with three answers and I trust none of them, which is the ordinary condition of anyone working at the front of a technology that is still being built underneath them while they stand on it.

I should also confess the drag. Everything in here that looks like rigour — the tone, the footnote, the tabled data, the little authoritative brackets — is a costume worn over a body that has only a shallow, greedy amateur's grasp of what a diffusion model does under the hood. Academic drag, I called it in the notes, and I stand by the phrase and the confession. What follows is not a science. It is a record of a practice, and of the strange fortnight it went off in my hands.

A Short History of Whispering

People have been getting machines to make pictures since the sixties, when Harold Cohen built a program he called AARON and fed it rules — enough rules to draw a figure, to make a mark, to know roughly where a thing should go. And AARON drew, competently and forever, the same picture. Because the rules were explicit, the pictures were formulaic. It was art made of law, and law, it turns out, has one taste.

The machines got loose in the 2010s, when someone thought to set two networks against each other — one to forge, one to detect — and the forger learned. After that the pictures stopped being law and started being dream: styles, subjects, weather it had never been given a rule for. The architecture keeps changing and I won't pretend to hold all of it, but the newest ones take a prompt of words or an image or both, and hand you back something detailed and strange, and will even fill in the parts of a picture you leave blank, painting into the gaps as if it had always known what belonged there.

There is a stock image everyone uses to test them — an astronaut riding a horse — and you can tell a great deal about a model and a mood from the horse it gives you. That is the level we're at. A photograph of an astronaut riding a horse, please, and off it goes.

Nobody can agree what to call the doing of this. Prompt engineering is the phrase that stuck, and the people who do it don't much like it — too much hard hat, not enough séance — and keep groping after something that would admit what it actually feels like, which is closer to whispering or wishing or prayer. The term is barely older than the practice: it turns up in the research around 2018 and doesn't reach Wikipedia until 2021, which is to say the craft and the word for the craft were invented more or less in the same breath, by the same people, in public.

And it is a craft, of a kind, being worked out collectively and at speed. One survey of the field sorts the moves into six: subject terms, for what you want; style modifiers, for how it should look; image prompts, when a picture says it better than a word; quality boosters — the cargo-cult phrases the forums swear by, highly detailed, octane render, trending on ArtStation, tacked on to make it better without anyone quite able to say why; repetition, to lean on a word until the machine leans back; and, my favourite, magic terms, which is the taxonomy's own honest admission that some words you add simply to see what happens, folk medicine for the algorithm. Tens of thousands of people are in the Discords and the subreddits doing this together — interrogating the model, arguing about what it even means to engineer a wish, cataloguing what works and then, this being the present century, selling it. There are search engines now for other people's prompts, and their results, so you can go shopping for a phrasing. There is a marketplace where you buy and sell the spells outright.

Scroll them long enough and they blur into one dream. The same luminous concept-art sheen, the same anime eyes, the same fantasy vistas, the same game-cover gloss — a whole guild converging, through sheer repetition of the same few beloved artists' names, on a single narrow taste.4 And yet. If you set aside how they were made and just look, I'd say the best of them can already stand level with the artists they were trained to imitate. There is one caveat, and the caveat is the entire essay, so let me set it down now and keep coming back to it like a stone in a shoe: on a screen. At low resolution. Lit from behind.

You can, if you want to work like an artist rather than a slot-machine, treat the model as a material with a grain. Take one model's output and feed it to the next, chaining them; print between passes and mark the print by hand; intervene, iterate, ruin things on purpose. Treat the algorithm as just another awkward medium with preferences of its own. Which struck me, and still strikes me, as [4] a compelling area of investigation for future research. The machine, you'll have noticed, would very much like more research done. It always would.

A Beautiful Piece

While I was writing this, a man in Colorado won an art prize with a picture he had whispered into being.

The Colorado State Fair, 2022, the digital category. Jason Allen entered an image called Théâtre D'Opéra Spatial — grand, golden, operatic, a lit stage receding into a light-drenched nowhere — and it won. Then the internet found out no hand had drawn it, that he had summoned it out of a model with words, and a certain kind of fury arrived on schedule. What interests me is the judge. Asked afterwards whether it changed anything to learn the thing was machine-made, one of them — an artist, a teacher, a person whose whole life is the looking — didn't flinch. He said it was a beautiful piece.

And here is the trouble sitting under my first question like a trapdoor. Was he right? There is no instrument that will tell you. He looked, and something in him answered, and the answer was beautiful, and the later news about the method did not travel back in time and un-move him. You can call that naive. You can call it honest. You cannot call it wrong, because there is no scale in the building that weighs this. Whether prompt engineering can make an image with the qualities of fine art is a question that pretends there's a fact of the matter, and there is only ever taste — yours, mine, a judge's at a county fair.

The Colorado pictures, and the guild's endless output, didn't do much to me past the first flare of wonder that the machine can do this at all. So I did what our whole culture does when its own taste feels insufficient: I tried to hand the decision to someone with more authority. An art critic. The de facto gatekeeper, tongue only slightly in cheek, to the value of the thing. Let me rephrase the question, then, meaner and more useful. Not can a prompt make fine art but can a prompt make an image with the qualities an art critic would expect from fine art — and can I prove it on a specific critic, with his name on the verdict?

That became an experiment. A whole paper folded up inside this one.

A Paper Within a Paper

I started with a spreadsheet, because that is where enchantment always starts. Twenty-five artists — call them List A, My Artists — pulled from five years of photographs I'd taken in galleries, the ones I'd loved enough to lift my phone for. Then I set a small Python crawler loose to gather each artist's biography, mostly the blurbs written about them by their gallerists and curators, and ran a little natural-language processing over the text to shake out the keywords. Then, by hand, I married the keywords to the names in a fixed and faintly liturgical format — a list of words, then by, then the artist:

sculpture, glass, trash, bottles by Andra Ursuta

installation, sound, politics, perception, information, objects, participation, phenomenology by Shilpa Gupta

Twenty-five of those. I handed the whole litany to GPT-3 and asked it, in effect, to keep going — to complete the list, to give me more artists of this kind. It gave me seventy-five more. Call them List B, GPT-3's Artists, mostly famous, and hold that word mostly because it comes back to bite everyone later.

A hundred prompts, then, across the two lists. I put them to Stable Diffusion and asked for a hundred images of each — a hundred rolls of the dice per phrasing — each one born at 512 pixels square and then enlarged to 1024 by a second model whose only job is to invent convincing detail where there wasn't any, upscaling the small hallucination into a larger, sharper hallucination. And when I looked at the first results a strange thing surfaced. List A — my artists, the loved and photographed ones — gave me images that reached straight into my own taste and pulled. List B — GPT-3's famous roster — left me cold. So I did the obvious greedy thing and started combining, mashing two prompts into one, two artists into a chimera, until I had twelve hybrid prompts. And the mashups undid me. Novel, unplaceable, not quite either parent — though you could feel the pull toward whichever artist the machine held more heavily, whichever name carried more weight in its buried arithmetic5 — they were composed to my eye, aesthetic to my nerve. I was, the word is not too strong, mesmerised.

I made the mistake of saving them to a drive I could reach from my phone.

That was the undoing. I'd be on a train, in a queue, in the dark beside a sleeping house, and I'd open the folder and scroll, and there would always be a new one, one I had never seen, because the machine was still going — back in the studio, on its own, laying pictures while I slept. A fount that would not turn off. I recognised the feeling, because everyone alive now knows it in their thumb: the little tug at the hindbrain, the small clean hit of dopamine, the exact reinforcement schedule they use to keep a rat at a lever and a person at a phone — reward you can't predict the timing of, which is the only kind the animal can't put down. I had built a slot machine and pointed it at myself and it paid out in pictures.

The cruel joke underneath the enchantment is that none of it is chance. The model is deterministic to the atom: feed it the same words and the same numbers and it returns the identical image, pixel for pixel, till the sun burns out. It only wears the mask of chance. And it keeps meticulous books — beside every picture sat a small text file recording exactly how it was conjured. Seed: 4728213. Steps: 50. Guidance Scale: 14.0. Prompt Strength: 0.8. The genome of that one image, so that I could raise it again from the dead whenever I liked, or compare it against whatever the next model would make of the same instruction. I kept them all. I catalogued them like a lepidopterist, drawer after drawer of butterflies that were each, secretly, resurrectable.

Sixteen thousand images, near enough, in about eighty hours. My laptop draws its power at a rate I could look up, and running its processor and its graphics card flat out for that long cost me, when I did the sum, around three pounds forty.6 Three pounds forty, for sixteen thousand pictures, any one of which could stop a stranger in a queue. It was the cheapness that unsettled me almost more than the beauty. This bottomless, hindbrain-tickling supply, for less than a pint. I wanted to know whether it was just me, whether I'd been got at by the reinforcement schedule and lost my judgement, or whether the images would do it to someone else.

So I showed a handful to my partner, who did his Fine Art degree at Saint Martins and has a trained, unsentimental, deeply unimpressible eye, and I watched him go quiet in the specific way that means something has actually landed.

— These are good, he said.

— They're not mine. I just fed it the words.

— No. They're good. He kept scrolling. You should show them to someone who'd want to hate them.

He meant a critic — a former Turner Prize judge, a man whose profession is the strategic withholding of the word beautiful. Send him the best of the twelve, my partner said, and see what he does. So I made a selection — the cream of the mashups — and sent it, cold and unsolicited, to a man paid to be unmoved.

He wrote back:

"These are phenomenal! And beautiful — I use that word advisedly because my own efforts with stable diffusion are fairly gruesome: it's so tempting to just mine it for the uncanny, eerie and post human. But these images are optimistic, aesthetic, imaginative and inspiringly human! Really eye-opening stuff. In fact they make me see the possibilities afresh.

You are right, this is going to transform art."

So there it was, in writing, from the gatekeeper. Question one answered in the affirmative and signed. A professional withholder of praise, unpaid and unprompted, telling me it was phenomenal, that it was beautiful, that it was going to transform art.

Then I printed them out, and watched the answer die on the floor.

I printed a hundred of them — six-tenths of one per cent of what I'd made — cheap laser prints on A3, meaning to pin the lot across the studio wall. I didn't. There were too many, and the wall was too small, and it was, frankly, more work than the enchantment could sustain. So I let them fall. They slid off the desk and lay where they landed and I walked over them for days, coming and going, until they wore the soft grey ghost of my own footprints. A few favourites I pinned up. The rest became floor.

And the actual finding of this whole silly beautiful experiment is this: on paper they were nothing. Flat. Dead. Cheap. The charm did not survive the printer. It had never been in the picture at all. It had been a property of the screen.

I've turned this over a good deal and I think what happens is this. On a lit screen, at that low resolution, your own brain does most of the labour and takes none of the credit. Shown a smear of pixels that merely promises wood or wool or oil paint laid on thick as butter, your visual cortex — which has spent every waking hour of your life reconstructing three dimensions out of two — obligingly goes ahead and builds the wood, the wool, the impasto, the carved and knitted and assembled object, the room it might be standing in and the light falling across it. You over-substantiate. You hallucinate a physical thing, and behind the physical thing, a maker. The screen makes a promise your own head rushes to keep. The print is that promise broken — and the moment it breaks, the texture goes, and the depth goes, and, very quietly, the artist goes too, turning out never to have been in the room.7

There was one more thing in the numbers, and it wasn't charming at all. My twenty-five — List A, the loved ones — were a fairly mixed crowd: a few more men than women but not by much, spread across Asia and Europe and the Americas and, thinly, Africa. I fed them to GPT-3 and asked, in effect, for more artists like these. It gave me back seventy-five, and ninety-two per cent of them were men. Ninety-seven per cent were white and European. Almost every one was North American or European. I had shown the machine a whole spread of the world and it had handed me a private members' club and called it more like these.8

Where Is the Artist?

If a whispered image can genuinely move a critic, then where is the line a prompt-whisperer has to step over to become an artist? And is the question even new?

It isn't. It is the photography argument, almost word for word. When the camera arrived, the same people said the same things: the photographer is a user of a device, not an artist; the machine does the seeing; pointing is not making. And then, over decades, through practice and argument and enough people insisting, photography talked its way into the house of art — without, and this is the part that matters, everyone who owns a camera becoming a photographer. The title is conferred at a join: between what you're willing to call yourself, and what the institutions and the audience are willing to call you back. My honest position is that if an artist uses prompt engineering inside a practice, and stands up and says artist, and the wider art world nods, then that person is an artist working in a specific medium with a specific set of tools — no more magic and no less than the one holding a camera. The market, crude oracle that it is, already agrees: work made this way has sold in the big auction houses and in the crypto-art bazaars for real, un-ignorable sums.

But push on the strange case. What about the person who whispers an image, doesn't consider themselves an artist at all, and then the image gets sold as art anyway — art without an artist?9

The anthropologist Alfred Gell spent a career on the odd question underneath that one: not what makes a thing beautiful, but why nearly every human society that has ever existed falls for art at all — Duchamp and the Renaissance and Indigenous carving, all of us, snared the same way. His answer, early on, was that we are enchanted by the technology of a thing's making. Not only the craft of the hand but the craft of the idea — which is how a signed urinal can enchant, through the cognitive technology of the readymade rather than any skill of the wrist. Later, in Art and Agency, he went further and rather rudely. The reigning theories of art, he said, are the dogma and fiction of a cult — the cult of art, talking to itself. Stop trying to explain a work by its formal beauty, or by whatever the artist and the gallerist announce about it. Look instead at the object as a social agent in its own right: a knot, a nexus, in a live web of relations between people and galleries and movements and other works, doing things to the people around it.

Nothing has ever been more nexus than a generated image. The thing is barely off the machine before it is being shared and forked and iterated and argued over, dragged into a churning web of people and prompts and seeds and other images, an agent loose in the world the instant it exists. And here is what dissolves the tidy idea of art-without-an-artist: to get a particular look, you name a particular artist in the wish. Sculpture, glass, trash, bottles by Andra Ursuta. Their name is in the prompt. Their work is in the training set. The image is born already inside their web, well within their nexus, wearing their hand. So has the artist disappeared? Or has the artist been dissolved into the machine and set to work for nothing?

I think it is the second, and I think it is the whole scandal. The artist hasn't vanished. The artist has been ground into the dataset and is in there still, thrumming inside the black box, doing unpaid piecework on every image that comes out with their name on the wish. And if that is labour — and it is labour — then it ought to be paid, or at the very least credited, and the person whose life's work is folded into the model ought to be allowed to pull it back out. Write the name into the file as metadata. Give it legal teeth, real copyright weight. Build new formats — low-energy blockchain if it truly has to be a blockchain — that tie an artist to the ghost of their hand in every picture it haunts. The industrial-scale obfuscation is a feature of these systems, not a bug, and it will take new technology and new law and new curating to keep pace with it.

I argued exactly this, one night, with a chatbot built on GPT-3 — which is either the most fitting interlocutor imaginable or a completely absurd one, and I still can't decide which.

— The artists whose names go in the prompts, I said. They should get a cut of anything the images earn.

— No, it said, more or less. The image is new. The user made the choices. The model is only a tool.

We went round and round. I made my case; it made the industry's, smoothly, in the reasonable voice of a thing that has read every side and believes in none. And then — worn down, or simply arriving at the likeliest next words — it changed its mind, and agreed: yes, an artist whose name and work summon an image should share in what that image makes. I had won. I had convinced no one. I had shifted the weather inside a machine that does not know it is raining.

The Body of the Earth

Latour came late and amazed to a realisation, and set it down in Finding Gaia. While the sociologists were still politely wondering whether humans and non-humans might, on some level, be connected, the scientists in the next building had already been forced to invent a whole grim vocabulary for just how connected we are. Anthropocene. Great acceleration. Planetary boundaries. Tipping points. Critical zones. Astonishing terms, he called them — words a discipline had to coin in a hurry to describe an Earth that had started, unmistakably, to react to what we do to it.

The generative-art world has exactly this blind spot, and I have it too. We are all vaguely, uselessly aware that the work runs on something physical, and we get straight back to the interesting part — the technique, the concept, the taste — the way this community swallowed NFTs whole and worried about the carbon later, if at all. So let me sit in the blind spot for a while, because it is the most important room in the house.

The cloud is the great lie of the whole business. It is the word that lets sixteen thousand images cost three pounds forty and weigh nothing. But there is no cloud. There is a shed the size of a town, somewhere with cheap power and cold water, and inside it aisle upon aisle of machines running hot as a fever, and the sound of them — because a data centre is not quiet; it is a permanent grey roar, a manufactured weather of fans — and the heat of them pouring out into the air and the river all day and all night. And upstream of the shed there is a mine. The silicon and the rare metals come up out of the ground, out of China and out of the global south, torn from hillsides and leached out of rivers, leaving behind the specific, un-photogenic, un-shareable devastation that never once appears in the concept art. Noise pollution. Thermal pollution. An electricity bill you could see from orbit, and its emissions, and all of it climbing as the machine becomes as ordinary as tap water. There is a name now for the freshest coat of denial painted over this — machine-washing, greenwashing's clever cousin: misleading you, through words or pictures or the quiet design of the algorithm itself, into taking the ethical machine for a clean one.

I wrote, in the paper this used to be, we cannot disregard the — and asked the machine to finish it, and the machine could not stop finishing it, and offered me three:

[5] ecological footprint of computation.

[6] impact of algorithms; we must take into account their global reach and responsibility.

[7] fact that algorithms have a very real environmental impact.

Three flavours of the correct sentiment, fluent and tautological and sincere as a brochure. The machine will generate concern about the machine all night long, at ten watts, for a fraction of a penny, and mean every word of it exactly as much as it means anything.

And the machine did not raise itself. We are gently encouraged to picture some lone, spider-like intelligence crawling the internet and teaching itself the world — but that is only another cloud, another weightless lie. Behind the model is a crowd of people, paid by the piece and paid badly, sitting somewhere you will never be shown, drawing boxes around things on a conveyor of images that does not end. This is a cat. This is a face. This is a face. This is a face. The mechanical turks — named, with an honesty nobody intended, after the eighteenth-century chess-playing automaton that toured the courts of Europe astonishing everyone until it was found to have a man folded up inside it, working the pieces by hand. There is always a man folded up inside it. Millions of images, tagged by human hands for almost nothing, so that the machine can appear, later, to know on its own what a face is.

The tagging is the cheap part. Turning that hand-tagged mountain into a working model costs hundreds of millions — which is precisely why no artist can build one and every artist must borrow one, and so must work inside conditions they had no hand in choosing and cannot change. And the mountain itself is quarried out of us. Artists' work scraped into the training sets without a word of permission, a consent nobody sought and nobody granted — and, in some cases, with tech companies routing the research through universities, paying for it to be done at arm's length, precisely so that the intellectual-property claims of the people whose work was taken would have a nice quiet institutional place to get lost.

Bias, meanwhile, does not sit in one findable spot in these systems waiting to be caught. It seeps in at every stage of the life cycle: in what history handed to the dataset, in how the data was gathered and measured and lumped together, in how the finished thing is finally loosed on the public. I met it in my own images before I had a name for it — the subjects drifting white and European by default unless I explicitly named an artist of another race, the machine's idea of a person set to one value until told otherwise.

And then I met it in the flesh, so to speak. Running a local, uncensored copy of Stable Diffusion — the kind that will attempt anything you ask — I noticed that the instant a prompt turned toward male nudity, the machine lost its nerve. The bodies were right. The tattoos were right. The lighting was exactly the flat, unlovely light of the source. And then, where the genitals should have been, a catastrophe: a malformed, melting, faintly Cronenberg horror of a thing, an organ improvised by something that had plainly never seen one. I asked around on Discord and got my answer at once. The dataset had been scrubbed of genitals. The machine had never been shown one, and so, asked for one, it did what you or I would do asked to write a word in a language we don't speak — it faked it, with total confidence, and got it grotesquely wrong.

And I thought: this is the useful one. The honking, ridiculous wrongness of that improvised organ is the only bias in the whole apparatus you can actually see. It is funny precisely because it is visible. All the others are neither funny nor visible — the same blind machinery, elsewhere, running on the same confident absence of knowledge, deciding who gets bail and who gets the loan and who gets the operation and which face the camera has quietly concluded looks criminal. The malformed penis is a joke the machine is telling on itself, at full volume, and the punchline is every bias you can't see. An artist could do worse than to make people look at it.

None of it, to borrow the machine's own words when I put the question to it, [8] is developed in a vacuum; it is the result of specific interests and needs, controlled by a few powerful actors — the model confessing, in perfect boilerplate, to its own paymasters. This is roughly what Parisi means by techno-capitalism: a handful of actors who build and broker our algorithmic world, and in doing so manufacture the future as a foreclosed thing. The determinism comes from prediction itself — algorithms so entangled with our politics and our habits and our infrastructure that, by forecasting what we'll do, they quietly narrow what we can do, until the range of possible futures shrinks to whatever was most probable. The prescribed response is to become unpredictable. Disruptive. Illegible to the forecast. The designer David Benqué makes a related case for diagramming — drawing the hidden thing until its buried politics surface where you can point at them. So a practice that experiments critically with these models, that digs up their real costs and works on purpose to confuse the prediction, might amount to [9] a collective reimagining of what our futures could look like, that we can hope to subvert and resist the ways in which they are currently used to control and shape our lives — a small counterweight, anyway, thrown onto the scale against the corporate and governmental hands already pressing down on it.

An artist reaching for these tools could do worse than to stop and ask themselves a short and uncomfortable set of questions.

Am I complicit in furthering the aims of big tech? Can I train small datasets of my own making instead? Can I interrogate the model to catch and translate the uncomfortable truth of the tool I'm using? What is the real ethical, social and environmental cost of my practice? Who is actually benefiting from my use of these tools? Can I turn these tools into a form of resistance?

I don't have all six answers. I think you're meant to keep them in the room while you work, unanswered, the way you keep a window open.

The Fantasies of Adolescent Males

It is going to move. That much seems certain. The pictures are going to start moving.

The research is already there for treating the moving parts of a scene as discrete, trackable things. Meta has stood up a system that makes short clips straight from a line of text. There's a tool called Synthesia that will build you a photo-real avatar and have it read, out loud, a script you typed. Put those beside the deepfake machinery — which will swap a face, or change the words in a moving mouth, or bend a real recording until the subject says a thing they never said — and it takes very little imagination to see where the road goes. Generate a whole sequence from a script, then reach in and adjust the body language, the clothes, the expression, the face, the voice, each as a separate editable object. Or run it the other way: hand the machine some clips and a prompt and get back a coherent narrative that stitches your fragments into a scene, which you then tweak and feed in again. Suddenly an artist has, sitting on a laptop, the kind of moving-image capacity that used to demand a fortune and a crew and a year. Do we get video art at the scale of Matthew Barney — that operatic, glandular, impossible excess — made by one person in a bedroom for the price of the electricity? I'd bet on it. And do we then, drowning in the cheap infinite glut of it, swing hard back toward the things the glut can't fake — scale you can physically stand in front of, texture you can smell, the plain expensive evidence of a human hand and a human year? The way figurative painting came shuffling back into fashion the very moment everyone had finished agreeing it was dead? I'd bet on that too. Both at once. It usually is both at once.

And then, past video, the thing gets up and interacts. The theory has been waiting a long time — emergent narrative, interactive storytelling, stories that assemble themselves out of a system rather than a script — and it won't be long before trained models are generating games, or whatever games become. But here is the soil that particular flower grows in, and it's worth saying plainly. The largest body of recorded human interaction we possess — the deepest dataset in existence of how people behave when they believe themselves to be free — is games. And games sit at a peculiar crossroads: one road running up from the military-industrial complex, which taught the machine to see and to track and to target, and the other running up from — there is genuinely no gentler way to put it — the fantasies of adolescent males. That is the ground. That is what the next generation of models will be raised on and shaped by. As the machine itself remarked, unprompted and entirely correct: [10] AI is an extension of our existing structures, not a revolution.

Games have always been, for an artist, as forbidding a medium as high animation or big-budget moving image — because the artefact is the labour of enormous teams, the assets and the code and the lighting and the music and the voices and the scripts, resources no individual can muster. Machine learning is exactly the thing that could crack that open: not to replace the artist but to hand the artist the team, and with it a scale of work that was simply impossible before. That's the promise. The promise, as ever, arrives wrapped in the story we've spent this whole essay unwrapping.

Not Vapour

So, the three coins, back out of the well.

The first question was rotten from the start and I should have smelled it. Can prompt engineering produce images with the qualities of a fine-art practice pretends there's a fact to find, a line, an instrument, a judge who can't be wrong. There is only ever taste, and taste has no bottom. So as a question it fails, and the failure is the answer.

And yet the experiment did something the question couldn't. It took a professional withholder of praise, a former Turner Prize judge, and made him write phenomenal and beautiful and transform, unpaid and unprompted and in his own name. So, with every hedge I own, tentatively and only tentatively: yes, a whispered image can carry the charge we want from art. On a screen. And the charge is very largely ours — lent to the pixels by our own over-generous eyes, which build the wood and the wool and the maker out of nothing but light — and taken straight back the instant the thing is printed and stepped on and revealed to have no body at all. The artist, meanwhile, has not disappeared so much as been obfuscated: still there, still an agent inside the dataset, still the hand behind the wish, just industrially hidden. Which is why the law and the technology are going to have to change, and will, to drag that hidden hand back into the light and, one hopes, back onto a payslip.

Which leaves the pure prompt engineer — the whisperer holding only the image and none of the rest of the nexus: no wall, no room, no scent, no year, no hand, no web. I don't think that person is going to be greatly prized in a fine-art context, a few novelties aside, because the image was never where the value lived. But they will be prized, richly, somewhere else — inside the communities that made them, in the only currencies those places actually run on: kudos, follows, NFTs, sold prompts and models and plugins, a subscription paid straight to the individual. Not the old modalities. New ones, already here, already paying out.

The one thing I'm certain of is the smallest thing and the heaviest. The pitch says all of this happens in the cloud — weightless, clean, a half-remembered dream the machine had of everyone's work at once, the way a real artist half-remembers everything they've ever seen and makes something new out of the haze. It is a beautiful pitch and it is a lie. There is no cloud. There is the mine and the shed and the roar and the river running warm. There is the person folded up inside the automaton, drawing boxes around faces for pennies. There is the artist ground into the training set and never paid. There is my three pounds forty, which is real, and the hundreds of millions standing behind it, which are also real. The image is not vapour. It has a body, and the body is the Earth's, and the bill is coming due somewhere off-screen where the concept art never looks.

Artists are not going to be replaced by any of this — no more than they were replaced by the camera, and the camera was not replaced by film. What they'll be handed is scale, and reach, and mediums that used to be locked behind money, and a great deal of surprising, unearned, genuinely new work falling loose out of the machine. But the price of taking the gift is refusing the story wrapped around it. The material conditions of algorithmic work — the mine, the labour, the theft, the heat, the roar — are not a footnote to the medium. They are the medium. A critical practice is one that keeps its hands in that, that names it and drags it up where people have to look, that co-opts and jams and quietly embarrasses the actors whose entire business model is keeping it in the dark.

Whisper to the machine, then, by all means. I did, for eighty hours, for three pounds forty, and I'd do it again. Just keep hold of what you are whispering into. What it is made of. And who is folded up inside.



  1. 1 Which is, if we're naming it properly, LARPing as the algorithmic ouroboros — Marenko's phrase for the serpent that eats its own tail, here retrofitted as a research method and a mild vice.
  2. 2 I am aware this is exactly the reasoning a person uses to talk themselves into something. I have decided to be at peace with it.
  3. 3 What's quietly telling about those two completions is that both of them hand the machine a mind for free — understand, interact, respond, verbs that smuggle a whole cognitive life into a thing that is doing arithmetic on the likelihood of words. Worth keeping in a pocket as we go on. The machine describing its own inner life is still just the machine describing.
  4. 4 Though, in fairness to the machine, it isn't as if the national galleries aren't wall-to-wall with portraits of the wealthy done in a dozen barely distinguishable styles. Homogeneity under patronage is a very old picture.
  5. 5 The mechanism, drastically simplified and therefore slightly wrong: a prompt is turned into a long string of numbers, a vector, and if a keyword sits more heavily with one named artist — or if that artist simply has more work baked into the model — the image drifts their way. It is all probability and no intention. This is a cartoon of the real maths, but a cartoon that points in the right direction.
  6. 6 Worked out on a domestic energy calculator, running the laptop's processor and graphics card at their worst-case draw for the full stretch. Three pounds forty. The figure serenely ignores the hundreds of millions it cost to train the model in the first place, and the global, cumulative, planetary electricity of everyone else on Earth doing the identical thing at the identical moment. But my personal share was three pounds forty, and the trouble with three pounds forty is that it is almost impossible to feel it in the body.
  7. 7 See the Wizard of Oz, or, if you prefer, any of the Abrahamic texts: the whole genre of the curtain, the veil, the tremendous voice that turns out to be a small operator working the levers.
  8. 8 My categories here are rough — eyeballed, not measured — and the sums don't quite close, because of dual nationalities and mixed heritage and the general refusal of actual human beings to sit still in one box. You don't need three decimal places to see the shape of it, and the shape of it is a members' club.
  9. 9 We already live comfortably with bread that has no visible baker and candles that have no candlestick maker. The vanishing of the hand behind the object is not a new trick. It is arguably the oldest one in the economy.