Introduction
Generative art, where images are produced through algorithms, positions the artist not only as a curator but as an orchestrator, establishing a collaborative bond between the artist and an elusive semi-autonomous system and issuing in a joint artistic endeavour (Pearson, 2011). Within a machine learning (ML) context, algorithms trained on features drawn from vast bodies of existing images can be made to produce pictures in the style of particular artists, styles or movements from algorithmic, written or visual prompts (Mazzone and Elgammal, 2019). To coax these models into producing an image a user supplies text prompts that set in motion a chain of deterministic and probabilistic processes, a practice that has come to be called prompt engineering (PE) (Oppenlaender, 2022).
My practice is one that explores the wider implications of algorithms and ML, and my intention here is to deepen that practice by attending to its very conditions — as Aldouri (2015) puts it, the “art historical legacies, particular institutional power … general institutional forms … and the material conditions of everyday life at a given moment”. I want to be precise from the outset about the position from which I write. I am one located artist: with a particular studio, a particular laptop, a partner who is a Central Saint Martins Fine Art alumnus and, as it turns out, a line of email to a former Turner Prize judge. The knowledge I can produce here is partial, embodied and interested. Rather than treat that partiality as an embarrassment to be disclaimed, I want to treat it as the ground of whatever the paper can honestly claim. What convinces me, what enchants me, what I can afford and who will answer my emails are not noise around the findings; they are the findings.
As an artefact of that practice, I propose that this paper is itself an art object, in that its construction enacts the same processes and methods as the other objects I orchestrate.
To deepen the practice I will learn from the communities that have formed around PE, experiment with popular generative text and image models, reflect on that experiment and then read the encounter back through the academic literature: the place of the artist in relation to PE; the ethical, environmental and social dimensions of ML that are too easily bracketed; and the likely futures of the practice. I organise this around three questions:
- Can prompt engineering produce images that have the qualities one would expect from a fine art practice?
- What theoretical concerns arise from using ML as part of a critical art practice?
- What opportunities will there be for artists with regard to prompt engineering?
Approach
This paper weaves human-machine encounters into its own form. It does so through three ML models; through experiments that respond to the theoretical concerns; and by feeding the output of one model back into my activity and my writing1. I define the approach and methods after the fact, following experimentation rather than before it. The linear format of the paper does not reflect the non-linear way it was made; the structure is a formal construct. Its purpose is threefold: to show the reader the shape a human-machine encounter can take inside a research paper; to develop my own intuition for these models as materials; and to encode, in the very structure of the paper, the problems that large language models pose to research2.
The models are GPT-33, Stable Diffusion4 and DALL-E 25. I chose them because they are available as commercial products with accessible APIs, and because each has drawn significant investment from big tech and finance (Hao, 2020; Krishna, 2022), producing large and active user communities and, less innocently, a dependence on the power structures that constitute techno-capitalism (Parisi, 2019). That choice is worth owning rather than naturalising: to work with these tools at all is to enter an enclosure someone else has already built and to accept its terms. There is no clean outside from which to make this work. I am using instruments I did not design and cannot fully see into, and the honest question is not whether the practice is compromised but what a compromised practice can nonetheless surface.
GPT-3 is an autoregressive language model that produces human-like text (Brown et al., 2020). I use it to generate fragments based on my own writing and on the paper’s references, as a form of automatic writing — a dream state (Bauduin, 2015) outsourced to a machine. This
[1] text will be used to explore how humans and machines can interact and understand each other
[2] response will be used to explore how machines can understand and respond to the text generated by humans6
will appear in a distinct typeface, referenced by a number in brackets, with each new line offering one of the options the algorithm generated; where GPT-3 answers with a number in square brackets I have changed it to round brackets for readability. The intention is to show the reader the texture of the generated text, to develop my own sense of how the model behaves and to stage, within the structure of the paper, the algorithmic ouroboros that Marenko names (Marenko, 2020) — the machine fed on its own and our outputs, doubling back on itself.
DALL-E 2 and Stable Diffusion generate novel images from text and image prompts (Ramesh et al., 2022; Rombach et al., 2021) and are used to respond to the paper’s theoretical concerns and to
[3] provide a visualisation of the potential interactions between humans and machines.
DALL-E 2 was chosen as the newest and most capable model available; Stable Diffusion because it is open-source and therefore more accessible — and, it turns out, more revealing, since an open model lets a practitioner reach the seams and blind-spots a closed one keeps hidden.
I take Merleau-Ponty’s good dialectic as a methodological caution: the arguments I venture are idealisations bound by language, and the gaps between “statements, thesis, antithesis and synthesis” (Merleau-Ponty, Lefort and Lingis, 1992, p.96) are vast and obscured by the very words that bound them. To bridge that gap I work through creative experimentation, as research through art (Earnshaw et al., 2015), treating the models as materials and letting method arise in response to theory, with personal reflection kept as an auto-ethnographic record. Which returns me to the paper’s performance of authority. I will still stage a certain amount of academic drag — the citations, the tables, the measured register — but I want to be clear about what the costume is doing. It is not a disguise meant to pass off shallow understanding as expertise; it is a way of making visible that all knowledge here is dressed, positioned and staged, mine included. The drag foregrounds rather than hides the situated, partial and interested character of what follows. I know these domains rudimentarily and from one angle. Owning that is not false modesty; it is the epistemic condition of the work.
Many theoretical concerns surfaced in an initial review of the literature; I chose a smaller subset (Table 1). This list does not capture the full breadth of what arises from PE but the excluded concerns were very much in mind throughout.
| Theoretical Concern | Chosen? | Reason for choice to include or not |
|---|---|---|
| Environmental Impact | ✓ | Mostly a footnote in most ML artist’s practices |
| Biases and datasets | ✓ | Without understanding this, a practitioner would have a large gap in their practice |
| Choice of model | ✓ | An integral part of using ML models |
| From images to video | ✓ | A personal interest with regard to art practice. |
| Gaming, emergent narrative and interaction | ✓ | A personal interest with regard to art practice. |
| Art and agency and where the artists sits as a practitioner when using ML | ✓ | Generative work forces us to consider novel ways of thinking about artists and artworks. |
| Techno-capitalism/determinism | ✓ | Actors within big-tech and finance are large forces within ML model distribution and gain the largest benefit from their use. |
| Art market and NFTs | ✕ | Theory and research is still quite nascent |
| Reifying generative work as physical artworks | ✕ | A fascinating concern that would warrant its own in-depth research and exploration |
| Performance and narrative agency | ✕ | A fascinating concern that would warrant its own in-depth research and exploration |
| The disappearing artist | ✕ | A bit too theoretically complex and applies to many other forms of art. |
| Making kin with algorithms | ✕ | A fascinating concern that would warrant its own in-depth research and exploration. |
Table 1: The theoretical concerns and reasons for their inclusion in this research paper
One row of that table deserves flagging now, because much of what follows argues against it. Against “Environmental Impact” I wrote that it is “mostly a footnote in most ML artist’s practices”. That is an accurate description of the field and a diagnosis of its failure. I have chosen to refuse the footnote. The material and ecological cost of computation is not a coda to this paper; it is one of its through-lines, and I will keep returning it from the margin to the centre where it belongs.
Prompt Engineering
Algorithms have produced images since the 1960s. The earliest example is AARON, built by the artist Harold Cohen from rules he defined by hand; because those rules were explicit, AARON produced formulaic images with a consistent style and figurative subject (Poltronieri, 2022). With the invention of generative adversarial networks (GANs) in the 2010s and successive innovations in ML, tools emerged that could generate images across many styles and subjects (Fig 1). The architectures differ but the most recent models can build increasingly sophisticated images from prompts made of text, images or both, and can even fill gaps in an image or add elements (Table 2).
Figure 1: Excerpt showing the timeline and output of the current generation of ML models, taken from Cetinic and She’s paper Understanding and Creating Art with AI: Review and Outlook (2021)
| Model Name | Type of Model | Type |
|---|---|---|
| Deep Dream7 | CNN | Text 2 dream, style transfer |
| Big sleep8 | GAN/CLIP | Text to image |
| GANBreeder9 | GAN | Combining existing images to breed more |
| DALL-E 210 | GPT-3->CLIP | Text to image, image & text to image, erase and replace |
| Midjourney11 | CLIP | Text to image |
| Stable Diffusion12 | CLIP | Text to image, image & text to image, erase and replace |
| Runway ML13 | Multiple | Toolset for text to image, image to image, text to 3d texture |
Table 2: A Summary of Popular Prompt-based Generative ML Art Tools
To use these tools you supply text, an image or both as a prompt, configure parameters that shape the computation and submit it to be processed into one or more images (Fig 2). The choice of model, parameters and prompt has material, aesthetic and procedural consequences. You can take the output of one model, alter it digitally or in print and feed it back in; by stacking models — GAN-chaining (Plain et al., 2022) — or passing one model’s product to another, you can iterate and transform, following instinct as you would with physical materials, with modifications and visual interventions between each iteration that an algorithm alone would not produce. Treating models as materials, with preferences specific to a practice, seems like
[4] a compelling area of investigation for future research.
Figure 2: Images produced from the prompt “a photograph of an astronaut riding a horse” with different parameters. This is often the default prompt used to compare different SD models and architectures.
These models are new, and the online communities around them are still working out the syntax and semantics of their activity. The accepted term, prompt engineering, came into active use after the release of GPT-3 and first appears around 201914; not everyone is comfortable with it, and some in the design and ML community want a word that better honours the emergent craft of coaxing valued images from a model.
| Modifier | Impact |
|---|---|
| Subject terms | The desired subject |
| Style modifiers | The style of the image |
| Image prompts | A visual prompt for subject and style |
| Quality boosters | Terms that increase the quality of the output |
| Repetition | Strengthen the association between terms |
| Magic terms | Terms that introduce unpredictable |
Table 3: A summary of different modifiers and their impact summarised from A Taxonomy of Prompt Modifiers for Text-To-Image Generation (2022)
Methods here are built collectively. In A Taxonomy of Prompt Modifiers for Text-To-Image Generation (2022) Oppenlaender identifies six modifiers that steer the generated image (Table 3). Communities of tens of thousands share prompts, building up a common repertoire of what produces images they admire (Oppenlaender, 2022). There are search engines for prompts and their images, such as Lexica15, and even a marketplace where prompts are bought and sold16. People interrogate the models, argue over what it means to engineer a prompt17, classify what does and does not work, then share and capitalise on the result.
Reviewing Discord servers18 and Reddit communities19, I would venture that the images these communities produce have a marked homogeneity — the same artists and prompts repeated, converging on the aesthetics of computer games, comics, anime and fantasy art (Fig 3). It is worth naming this homogeneity precisely, because it is not a quirk of taste but a structural tendency. A few high-yield names and modifiers reliably return admired images, so they are replanted endlessly, crowding out the varieties that do not. This is, transposed from soil to latent space, what Vandana Shiva called a monoculture of the mind: a diverse commons of visual culture narrowed to a handful of profitable cultivars, its resilience and its strangeness bred out in favour of predictable yield. And like agricultural monoculture it rests on an enclosure — the uncredited harvesting of countless artists’ work into a seed-bank no one consented to plant. Despite this narrowing, the scale of collective experimentation has produced genuinely sophisticated images, and I would argue that, method aside, they stand visually alongside the work of the artists the model was trained on, with one caveat: on a screen, at low resolution.
Figure 3: Generative images typical of those being produced. Sourced from Lexica.art. The images have no copyright.
Figure 5: Théâtre D'opéra Spatial (Allen, 2022)
But is it Art?
While I was writing this paper an image made with PE (Fig 5) won a prize at the Colorado State Fair, in the digital arts / digitally-manipulated photography category (Roose, 2022); one of the judges, an artist and art teacher, on learning the image was made with ML was untroubled and called it a “beautiful piece” (Metz, 2022). On the strength of that award, could you as a reader answer the first question — whether PE can produce images with the qualities of a fine art practice? I suspect your answer depends heavily on your own position, taste and sense of what a fine art practice is. Personally these images, and those in the PE communities, do not move me past an initial wonder that a machine can make them. So might we hand the decision to someone in a position of cultural authority — an art critic? The suggestion is only half tongue-in-cheek: critics function as de facto gatekeepers of art’s value. But it is worth remembering that a critic is a located authority too, formed inside particular, largely Western and market-linked institutions of taste. To convince one is to learn something about the situation of the critic and the screen, not to settle the nature of the image. Let us ask, then: can PE produce images with the qualities an art critic would expect from a fine art practice?
Can we Convince the Critic? An Experiment. Or, A Paper Within a Paper.
Methodology
The experiment began with a spreadsheet of 25 artists (List A: My Artists), selected from my own photographs taken during gallery visits over the past five years. A Python web crawler gathered text for each artist and Natural Language Processing (NLP) extracted keywords from their biographies — biographies predominantly written by gallerists or curators on gallery or personal webpages. That provenance matters more than it first appears: the keywords that seed the entire pipeline are not raw description but the language of the market’s intermediaries, the terms in which an institution has already agreed to sell each artist. The extraction begins before the model does.
The keywords were then combined by hand with the artists’ names into prompts of the form “list of keywords, by artist”:
sculpture, glass, trash, bottles by Andra Ursuta
installation, sound, politics, perception, information, objects definitions, participation, inter-subjectivity, phenomenology by Shilpa Gupta
This list of my artists was used as a prompt to have GPT-3 complete a further list of 75 artists (List B: GPT-3 Artists) through its sentence-completion.
Stable Diffusion then produced 100 images for each of the 100 prompts drawn from Lists A and B. Each image, initially 512x512 pixels, was upscaled to 1024x1024 with the Real-ESRGAN model.
Commentary, Observations and Insights
The first thing the images revealed was a divergence in appeal. Those from List A had an aesthetic that aligned with my taste; those from List B, mostly well-known artists, did not. This led me to combine prompts from different artists, producing 12 hybrid prompts.
Reviewing the mashups I was mesmerised. Novel relative to any single artist, they reflected my taste in composition and aesthetic; as hybrids they were not identifiable as any one artist, though depending on the combination the image could skew toward one — presumably because that artist carried heavier weightings for the prompt’s words20.
I set the images to save to a cloud drive so I could review them away from the studio, and became addicted to endless scrolling, eventually enchanted by the images from the 12 hybrid prompts. The pleasure was exactly that of the variable-interval reinforcement that underpins gamified social media and gambling machines. It is worth pausing on what that means, because it inverts the usual story of who is training whom. In that moment the machine was not my instrument; my attention was its product. The scroll is a designed capture of the body, tuned to the reward circuitry rather than to judgement, and it works by reducing an open, reflective encounter to compulsion. A critical practice that hopes to resist reduction has to begin by noticing that the reduction is being performed on the practitioner, in the pleasure centres, in real time.
The diffusion model is deterministic: the same prompt and parameters always return the same image, through complex probabilistic computation. My tool saved, alongside each image, a .txt file recording the prompt, parameters and model, so I kept a record of the conditions that produced each image, to reuse and to compare against future architectures.
In total the experiment produced roughly 16,000 images over about 80 hours. My laptop draws 0.125 kW/h; I used around 10 watts of energy costing me about £3.4021 — a trivial sum for 16,000 images (Fig 7). That trivial figure is the most ideologically loaded number in the paper, and it is worth refusing to take at face value. £3.40 is the marginal cost to me. It is cheap precisely because the expensive parts were paid elsewhere and earlier by others: the hundreds of millions spent training the model, the electricity and heat of the data centres, the miners whose landscapes were opened to make the silicon, the labellers who tagged the data for pennies and the artists whose work was harvested without consent to constitute the latent space. £3.40 is not the cost of the images; it is what enclosure feels like from the inside, once the costs have been successfully externalised onto people and places kept out of view. The cheapness is the tell.
Producing 16,000 images this cheaply, this endless fount tickling my hindbrain with dopamine, was curious and unexpected, and I wanted to know whether it was only me. I shared a sample with my partner, a CSM Fine Art alumnus; similarly enchanted, he suggested I compile a selection22 for the art critic and former Turner Prize judge, who replied:
“These are phenomenal! And beautiful - I use that word advisedly because my own efforts with stable diffusion are fairly gruesome: it’s so tempting to just mine it for the uncanny, eerie and post human. But these images are optimistic, aesthetic, imaginative and inspiringly human! Really eye-opening stuff. In fact they make me see the possibilities afresh.
You are right, this is going to transform art.”
To interrogate my own enchantment I printed 100 of the images onto A3 paper and also ran them continuously on a studio screen. I had meant to pin all 100 to the wall but did not — too much work, too little wall, and in any case only 0.6% of the total produced. I let some fall to the floor, where they gathered footprints; a few favourites I pinned up. The printed images lost much of their charm: cheaply laser-printed and, more tellingly, stripped of the texture a screen had seemed to promise. They became flat and less interesting.
Why? If we look at Figure 8, the architecture of the brain is transforming a 2D image into a sculpture made from wood and wool. But the textures, the lighting, the implied spaces and objects are emergent properties of a dataset and an algorithm; because we are so practised at reading 2D representations of 3D things, our own cognitive apparatus generates real texture, physical sculpture and drawn and painted surfaces out of the image. We over-substantiate — we promise, on the evidence of a probabilistic distribution of pixels on a screen, extant places and objects that were never there. Once printed, an image whose promise was a painting, drawing, collage or carving breaks that promise; the flat sheet exposes the gap between the image on the screen and its reality as a digital artefact.
Figure 8: The image produced using the prompt 19. news, history, gossip, oil, painting, lubugo, cloth, narrative, storytelling, bodies, figures, dystopia, sculpture, everyday, assemblage, wood, knitted, natural, craft by Michael Armitage and by Alexandra Bircken and variables Width: 512 Height: 512 Seed: 4728213 Steps: 50 Guidance Scale: 14.0 Prompt Strength: 0.8 Use Face Correction: GFPGANv1.3 Use Upscaling: RealESRGAN_x4plus Sampler: euler_a Negative Prompt: Stable Diffusion Model: C:\stable-diffusion-ui\stable-diffusion\sd-v1-4.ckpt
Reflections on Diversity and Representation
Comparing the two lists (Table 4) exposed a pronounced bias in List B, the list GPT-3 generated, across gender, nationality and race.
| Artists’ Characteristic | Distribution List A: My Artists | Distribution List B: GPT-3 Artists |
|---|---|---|
| Gender | 56% Male 44% Female | 92% Male 8% Female |
| Nationality | Asia: 32% Europe: 28% North America: 28% Africa: 4% South America: 4% | 59.09% North America 40.91% Europe 2.27% Asia |
| Race | 33.33% Asian 29.17% White European 20.83% Black 12.5% Latin | White: 97.37% Black: 1.32% Asian: 1.32% |
Table 4: Distribution of artists by gender, nationality and race. The LLM output had a clear male & eurocentric bias despite the distribution of the artists in the prompt that was used.24
The usual way to read this table is as three separate failures — a gender bias, a race bias and a geographic bias — to be totted up and deplored. I think that reading is too weak, in two specific ways.
First, the categories themselves are not neutral instruments. To sort artists into male or female, into White, Black, Asian or Latin, is already to apply a grid that was drawn under colonial conditions and that does not fit the bodies it is laid over. My own note to the table concedes as much: the numbers are “slightly off” because of dual nationality, mixed heritage and ambiguity, so that some artists fall into two boxes at once24. That slippage is not a rounding error to apologise for; it is the evidence. The bodies exceed the boxes. What the pipeline does is take a messy, plural world of practising artists, force it back through classificatory categories that colonial administration invented and then amplify them — returning the grid to us hardened and cleaned of its exceptions.
Second, the three biases are not additive but compounding and co-constituted. GPT-3 does not merely prefer men, and separately prefer white artists, and separately prefer North America and Europe. It reconstitutes a single figure — the white, Euro-American, male artist — as the default meaning of the word “artist”, against which everyone else registers as a marked deviation. The clearest evidence is in the images themselves: in my experiments the human subjects skewed white European unless I named an artist of another race in the prompt. Whiteness there is not one option among several; it is the unmarked ground that becomes visible only when something is placed against it. My own List A was already diverse — 56% male, and spread across Asian, White European, Black and Latin artists, across five continents. The system did not so much fail to reflect that diversity as actively regress toward the canonical mean, laundering a plural input back into the monoculture. The pipeline is a machine for restoring the centre.
Conclusion
The experiment loosely indicates that one can captivate both the orchestrator and a fine-art audience with the potential of PE — but only on a screen. Printed, the images lose the material qualities we confabulate from a screen on the assumption that we are looking at a photograph of a physical work. On a screen we assume a physical artefact and therefore a maker; in print the illusion shows itself, the promise breaks and both the artist and the art object disappear25.
Where is the Artist?
If we can make images that convince a critic of PE’s potential, where is the boundary a prompt engineer must cross to become an artist? Twentieth-century art history has worried this question about photography, Dadaism, ephemeral work, abstract expressionism, video art, performance art, land art and more; is prompting an algorithm any different? The same charge was once laid against photography — that the photographer merely operated a machine that did the work. Over time the language and practice of photography developed and it became an accepted art, though not everyone who takes photographs is a photographer or an artist: the distinction is made at the interface between a photographer’s self-categorisation as an artist and its confirmation by institutions and audiences. So too, I would argue, with PE: an artist who uses prompt engineering, positions themselves as an artist and has this confirmed by the wider art community is an artist with a particular set of methods, materials and tools. If we take the market as our crude measure, ML works already sell for significant sums in galleries (Christie's, 2018) and in online cryptoart circles (Kent, 2022).
But what if the person prompting the image does not consider themselves an artist, and the work is later sold as art — art without an artist?26 In an early essay the anthropologist Alfred Gell tried to ground a theory of art anthropologically rather than aesthetically, semiotically or art-historically, reconciling Duchamp, Indigenous art and Renaissance painting to ask why nearly all human societies value art objects as they do. We are enchanted, he argued, by the technology of an object’s making — not only the physical craft but the cognitive technology of its ideation, hence Duchamp and the readymade (Gell and Hirsch, 1999). In Art and Agency (1998) Gell goes further: existing theories of art are the dogma and fiction of a cult, and he offers instead a model in which artworks are themselves social agents, existing as a nexus of social relations around a work (van Eck, n.d.). Rather than explain art by its “formal or aesthetic value” or by what artist or gallerist says of it, we should attend to the technologies, cognitive and practical, that enmesh a work in a network of humans, galleries, movements and other agents with their own relationships to the work (Gell, 1998).
That an artwork becomes an agent in its own right, beyond the artist, seems especially apt for prompt-engineered images, which circulate through networks of people who share, iterate and build on both the images and the conditions that produce them. But the nexus cuts the other way too. To make an image with particular qualities you name particular artists in the prompt, and those artists’ work sits inside the dataset the model was trained on; the generated image is squarely within the named artist’s nexus. So has the artist really disappeared? I would argue the artist’s labour is present inside the algorithm and its training data, and that any artist named in a prompt — or whose work helped generate an image — remains an author of it. Here is where I part company with the ethics-footnote framing of this problem. The uncredited absorption of a named artist’s work into a model, harvested without consent and resold as a service, is not a lapse of etiquette to be smoothed over with an acknowledgements page. It is extraction — the enclosure of a commons of human creative labour, the same logic by which a shared inheritance is fenced, renamed as private stock and rented back to the people it was taken from. It is, in the older and more exact sense of the word, a biopiracy applied to culture.
If that is the right description, the remedies follow as questions of who captures the value rather than of good manners. Labour should return a share of the value it creates, or at the very least explicit credit for the image27, and the ability to withdraw from the dataset. Should we encode the artist’s name into these images as metadata, with legal and copyright protection, or design new file formats that use low-energy blockchain technologies to bind artists to the artefacts made from their work? The deeper point is distributive. The value these systems produce was constituted by the very people the systems write out; the platform climbs on a ladder built from their labour and then kicks it away beneath them. Perhaps the prompt engineer who declines the title of artist declines it precisely because they sense this — that the artist has not disappeared but persists, unpaid, thrumming inside the algorithmic black box. I expect we will see technological, curatorial and legal innovation racing to keep pace with the industrial-scale obfuscation on which the most popular models are built.
The Hidden Cost of Machine Learning
In Finding Gaia (2017), Latour reflects on his surprise, as a sociologist, at the relationship between the human and the nonhuman:
“We were still discussing possible links between humans and nonhumans, while in the meantime scientists were inventing … ways to talk about the same thing … on a completely different scale: the “Anthropocene,” the “great acceleration,” “planetary limits,” “geohistory,” “tipping points,” “critical zones,” all these astonishing terms …. terms that scientists had to invent in their attempt to understand this Earth that seems to react to our actions” (Latour, 2017)
That blindspot — that the human and nonhuman are intrinsically intertwined — describes the generative art community too, where most practitioners are only vaguely aware of the physical substrate their work depends on28 while absorbed in its technical, conceptual and aesthetic surface; witness how quickly the community adopted NFTs despite their environmental cost. Algorithms run on silicon, in devices whose physical parts are made through extensive mining that produces environmental devastation in China and across the global south. This geography is not incidental. It is the old colonial map redrawn: the extraction sited where its consequences can be borne by people far from the aesthetic pleasure, the harm exported to the periphery so the centre can enjoy a clean and weightless “cloud”. Onto this is layered a new practice of concealment — “machine-washing”, which “involves misleading information about ethical AI communicated or omitted via words, visuals, or the underlying algorithm of AI itself” (Seele and Schultz, 2022). Nor is it only mining and manufacture: the data centres where computation happens carry a staggering ecological cost — noise pollution, thermal pollution and vast electricity use (Suresh and Guttag, 2021; Monserrate, 2022), with the emissions that follow, and it will only grow as ML becomes ubiquitous. We cannot disregard the
[5] ecological footprint of computation.
[6] impact of algorithms, we must take into account their global reach and responsibility.
[7] fact that algorithms have a very real environmental impact.
and too few artists, in thrall to the surface of ML, are critical of its environmental cost. A critical machine learning practice does not flinch from this dimension; it actively seeks out and explores ways of working that reduce the harm.
It would help those who use these models to understand that they are not the product of some spider-like artificial intelligence crawling across the internet and reading images and text by itself. They require enormous human labour to categorise the data — people working as poorly paid “mechanical turks”, marking features on an unrelenting stream of images (Amoore, 2020; Suresh and Guttag, 2021). The name is more honest than those who use it intend. The original Mechanical Turk was a fraud: a chess-playing automaton with a human being concealed inside, labour disguised as machinery for an audience that preferred the magic. Today’s arrangement is the same trick at planetary scale — hidden human work, disproportionately sited in the global south and disproportionately racialised and gendered, dressed up as autonomous computation. This is time-consuming labour across millions of images, outside the scope of any individual artist (Monserrate, 2022), followed by the vast computational cost of training a model from the data (Incze, 2019), with hundreds of millions spent on the most successful models. Because that labour is beyond any artist, we must rely on models built by tech companies and research institutions — tweaking output with fine-tuning such as textual inversion in Stable Diffusion (Foong, 2022) — and so can only make work that rests on these fraught environmental and ethical conditions. And is it ethical when the tagged images were made by artists included in datasets without their consent, with tech companies paying universities to conduct the training research precisely so as to circumvent intellectual property claims from those whose work is used (Calma, 2021)? The academy is here conscripted to launder the extraction — to give enclosure the alibi of research.
The choice and construction of a dataset shapes a model’s interpretation and output, introducing historical, representation, measurement, learning, evaluation, aggregation and deployment biases before the model ever reaches the public (Fig 9) (Mehrabi et al., 2021). It is worth insisting that these are not one bias sampled at seven points but a bias compounded through a lifecycle, each stage inheriting and amplifying the last, which is exactly why it cannot be patched at the output with a fairness filter. This could be certain popular artists and movements favoured over others, skewing output, or specific races favoured in the production of images — something I met directly, where the subjects skewed white European unless an artist of another race was named. And during my experiments with SD I kept meeting malformed penises whenever a prompt shifted to male nudity, so I made the bias an object of study (Fig 10); on Discord I learned the training data included no genitals. Working directly with an open model, an artist can find the blind-spots and the implicit and explicit biases built into it. There is a biopolitics legible in that comic failure. The model was trained on desire — on scraped pornography among everything else — and then scrubbed of the organs of desire, so that it renders bodies fluent in every respect but the one that was governed out of the data. What a system can and cannot depict is a quiet politics of which bodies are legible, and the grotesque, visible failure stands in for the far more consequential failures that stay invisible — the biases, harder to discover, in models used for prison sentencing, stereotype reinforcement and financial and health discrimination. Artists should not only be aware of these biases; they can capture and communicate them through their practice.
Figure 9: A diagram taken from Suresh and Guttag’s paper showing where bias sits in the machine learning lifecycle (2021)
Fig 10: Blind spots in the dataset: using a pornographic image of a deceased porn star, with a clearly erect penis, and a pornography-related prompt, I prompted a local, uncensored build of Stable Diffusion. The lighting, body-shapes, tattoos and other elements match the source and what pornography leads us to expect; the penis alone is malformed, the gap in the dataset clearly articulated to a viewer.
In Dreams Rewired (Manu Luksch, Martin Reinhart, Thomas Tode), Tilda Swinton narrates a history of communication and technology in which the interests and needs of governments and corporations are shown to be built into how we use it. So when we use ML models we should remember that they are
[8] not developed in a vacuum, but are the result of specific interests and needs, as well as a result of a process that is controlled by a few powerful actors. ML is often presented as a scientific process with the goal of improving efficiency and productivity. However, algorithms are often designed to achieve a specific goal, and these goals are often determined by the interests of a few powerful actors.
Parisi argues that techno-capitalism — the actors who create and mediate our algorithmic world — is generating deterministic futures: algorithms, through their predictive power and their entanglement with our political, social and physical structures, define and reduce future possibilities, so that we must fight back by becoming unpredictable, disruptive and radical (Parisi, 2019). This is the hinge of the whole argument, so let me sharpen it. Determinism is not a property of the mathematics but a political project: it is the promise, sold as inevitability, that the future is already computed and that the only sane response is to optimise inside it. That is the deep grammar of technosolutionism — every problem, including the problems the technology itself creates, awaiting a further technical fix. To resist reduction is to refuse that grammar: to keep open the futures the prediction would close. In a talk, the artist and academic David Benqué (2020) argues for the design practice of diagramming as a way of interrogating algorithms and surfacing their hidden political and social dimensions. An art practice that experiments critically with ML, unearths its hidden costs and works deliberately to confuse algorithmic prediction could contribute to
[9] a collective reimagining of what our futures could look like, that we can hope to subvert and resist the ways in which they are currently used to control and shape our lives
opening new avenues of inquiry and helping to counterbalance the activity of corporate and government actors as we contend for non-dystopian futures.
Reflecting on the concerns we have covered, an artist using generative tools or ML may benefit from asking:
- Am I complicit in furthering the aims of big tech?
- Can I train small datasets that I create myself?
- Can I interrogate the models to capture or translate this uncomfortable truth of the tool I am using?
- What is the ethical, social and environmental impact of my practice?
- Who is benefiting from my use of these tools?
- Can I use these tools as a form of resistance?
The Future
Looking forward, it seems inevitable that these models will advance to generating moving images. There is already research treating moving elements as discrete, trackable units (Fan et al., 2021), and Meta has a text-to-video system29 that produces short clips from prompts; there is also Synthesia, which makes photorealistic avatars that speak a script you have written30. It is easy to imagine tools that produce extended sequences from text, image and video prompts — realising whole scripts, or turning existing footage into scenes with ML-modified mouth movements. Add the models behind deepfakes, which swap faces or change what a subject appears to say, and one can foresee tools that generate video from a script and then tweak body-language, clothing, speech, expression and faces within it; or, conversely, that take clips and a prompt and combine models to transform them into a consistent narrative to iterate on. Suddenly artists have access to video at scales that once demanded enormous money, time and resources. Will we see an explosion in the scale of video art — work that appears to be at the scale of Matthew Barney but is entirely generative? And, conversely, a renewed appreciation for scale, texture and craftsmanship, in the way figurative painting became popular again after a period of decline, abstraction and conceptual exploration (Cullen, 2019), while some creative jobs are lost en masse?
I read that renewed hunger for craft not as nostalgia but as a symptom worth taking seriously: a reaching for the material, the embodied and the situated that the screen cannot deliver, the same over-substantiation that broke in my print experiment surfacing as a collective appetite for the thing itself. It is a counter-current to the monoculture.
Following ML through text, image and video, I propose it will continue into generative gaming and interaction. The theoretical ground is laid — emergent narrative and interactive storytelling (Louchart and Aylett, 2004) — and it will not be long before trained models can generate games or interactive artefacts. But gaming is the largest dataset in the field of interaction, and it sits at what could be argued as the intersection between the military-industrial-complex and the fantasies of adolescent males and
[10] we must be aware that the current AI models are replicating and amplifying existing power structures and behaviours. AI is an extension of our existing structures, not a revolution.
That is precisely the risk to name: the largest interactive corpus encodes a particular structure of desire — militarised, masculinist — and a generative gaming built on it will not transcend that structure but amplify it, unless artists intervene in what is being modelled and for whom. As a medium, gaming has been inherently difficult for artists to work with, as complex animation and moving image are, because the artefacts demand huge teams of specialists and large resources to produce the assets, code, music, lighting, voice acting, scripts and other elements of gameplay. But ML could make this more accessible, letting artists include these technologies in their practice and work at otherwise impossible scales; the question a critical practice keeps asking is whose desires the accessible tools have already been trained to serve.
Conclusion
Given the pace and sophistication of prompt-based ML models and the speed with which PE has been adopted, we should expect continued advances and new methods to emerge, with artists and prompt engineers responding and evolving their practice amid an enormous volume of collaboration and experimentation from entrepreneurs, computer scientists and prompt engineers. Observing the communities and the early research into the taxonomy of prompts, we have seen an increasingly sophisticated practice forming around PE, even if its output is somewhat homogeneous31.
Reflecting on the images made by these communities, the subjectivity of deciding what is and is not fine art is itself a weakness in the first research question. But the experiment inspired and directed work with an ML model that showed images made with PE can excite an art critic, so I would tentatively venture that PE is likely to produce images with the qualities one expects of a fine art practice — with the crucial qualification that this happens on a screen, and that the qualities we grant these images are largely an artefact of our own cognitive apparatus, there being no real texture, scale or other properties that enchant us before physical art objects. Art does not only live on screens. It lives in galleries and on walls and in the hearts of an audience; it has scent and texture and other characteristics that a digitised image lacks. Reading art through Alfred Gell’s anthropology — art without an artist — we find the artist obfuscated but plainly present as an actor inside the training data, and we should expect legal and technological change to follow. But with only the image and none of the rest of the nexus, pure prompt engineers are unlikely to be much valued in a fine art context, a few novelties aside — though they will be valued and rewarded within PE communities, outside the standard fine art modalities32.
Contrary to the pitch from big tech, the work made through ML does not exist in the cloud as digital vapour, a half-remembered cloud of work that lives in the mind of a human artist and inspires them. It is the product of many physical processes with environmental, ethical and social costs, and this is the disconnect between signifier and signified that artists are equipped to expose. My own experiment is a small piece of evidence for exactly this: a set of images that could enchant a former Turner Prize judge on a screen and then, printed and dropped on the floor to gather footprints, give up the illusion entirely. The enchantment and its breaking are the same finding seen twice. Artists with a critical practice can break the enchantment deliberately — can explore and interrogate the material conditions of generative models, make the hidden costs felt, name the extraction as extraction and the enclosure as enclosure, stand with the dispossessed labour thrumming inside the machine and press the distributive question of who is paid and who is written out. Artists are not going to be replaced by machine learning, any more than they were replaced by photography, or photography by film. ML will instead give artists the potential to work at scales and in media that would otherwise be too expensive or difficult, generating many unexpected and surprising outcomes — but in doing so they must refuse the monoculture, contest the marketing machine of techno-capitalism and ensure that the material conditions of algorithmic work are explored, interrogated, challenged and counteracted.
Reference list
Aldouri, H. (2015). On critical art practice: some preliminary reflections. [online] Artblog. Available at: https://www.theartblog.org/2015/11/on-critical-art-practice-some-preliminary-reflections/.
Allen, J.M. (2022). Théâtre D’opéra Spatial. [2022] The New York Times. Available at: https://static01.nyt.com/images/2022/09/01/business/00roose-1/merlin_212276709_3104aef5-3dc4-4288-bb44-9e5624db0b37-superJumbo.jpg?quality=75&auto=webp.
Amoore, L. (2020). Cloud ethics: algorithms and the attributes of ourselves and others. Durham: Duke University Press.
Bauduin, T.M. (2015). The ‘Continuing Misfortune’ of Automatism in Early Surrealism. [online] ScholarWorks@UMass Amherst. Available at: https://scholarworks.umass.edu/cpo/vol4/iss1/10/ [Accessed 03 Oct. 2022].
Benqué, D. (2020). Hybrid Futures: Diagrams for Critical Algorithmic Practice – A talk by David Benqué. [online] www.youtube.com. Available at: https://www.youtube.com/watch?v=c7BFTOzc59U [Accessed 13 Nov. 2022].
Brown, T.B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D.M., Wu, J., Winter, C. and Hesse, C. (2020). Language models are few-shot learners. CoRR, [online] abs/2005.14165. Available at: https://arxiv.org/abs/2005.14165.
Calma, J. (2021). The Climate Controversy Swirling around NFTs. [online] The Verge. Available at: https://www.theverge.com/2021/3/15/22328203/nft-cryptoart-ethereum-blockchain-climate-change.
Cetinic, E. and She, J. (2021). Understanding and Creating Art with AI: Review and Outlook.
Christie's (2018). Is artificial intelligence set to become art’s next medium. [online] Christies.com. Available at: https://www.christies.com/features/A-collaboration-between-two-artists-one-human-one-a-machine-9332-1.aspx.
Cullen, M. (2019). Why has figurative painting become fashionable again? The Spectator. [online] 5 Sep. Available at: https://www.spectator.co.uk/article/why-has-figurative-painting-become-fashionable-again/ [Accessed 13 Nov. 2022].
Earnshaw, R., Liggett, S., Cunningham, S., Heald, K., Thompson, E. and Excell, P. (2015a). Models for research in art, design, and the creative industries. doi:10.1109/ITechA.2015.7317457.
Earnshaw, R., Liggett, S., Cunningham, S., Heald, K., Thompson, E. and Excell, P. (2015b). Models for research in art, design, and the creative industries. doi:10.1109/ITechA.2015.7317457.
Fan, H., Xiong, B., Mangalam, K., Li, Y., Yan, Z., Malik, J. and Feichtenhofer, C. (2021). Multiscale vision transformers. [online] doi:10.48550/ARXIV.2104.11227.
Foong, N.W. (2022). How to Fine-tune Stable Diffusion using Textual Inversion. [online] Medium. Available at: https://towardsdatascience.com/how-to-fine-tune-stable-diffusion-using-textual-inversion-b995d7ecc095 [Accessed 13 Nov. 2022].
Gell, A. (1998). Art and Agency. Clarendon Press.
Gell, A. and Hirsch, E. (1999). The Technology of Enchantment and the Enchantment of Technology. In: The Art of Anthropology: Essays and Diagrams (1st ed.). Routledge.
Hao, K. (2020). The messy, secretive reality behind OpenAI’s bid to save the world. [online] MIT Technology Review. Available at: https://www.technologyreview.com/2020/02/17/844721/ai-openai-moonshot-elon-musk-sam-altman-greg-brockman-messy-secretive-reality/.
Incze, R. (2019). The Cost of Machine Learning Projects. [online] Medium. Available at: https://medium.com/cognifeed/the-cost-of-machine-learning-projects-7ca3aea03a5c.
Kent, C. (2022). NFTs Can Be Artistically Groundbreaking — Meet the Artists and Curators Leading The Way. [online] ARTnews.com. Available at: https://www.artnews.com/list/art-news/artists/what-is-best-nft-art-1234631062/and-the-virtual-joins-the-physical-world/.
Krishna, S. (2022). Stable Diffusion creator Stability AI accelerates open-source AI, raises $101M. [online] VentureBeat. Available at: https://venturebeat.com/ai/stable-diffusion-creator-stability-ai-raises-101m-funding-to-accelerate-open-source-ai/ [Accessed 13 Nov. 2022].
Louchart, S. and Aylett, R. (2004). Narrative theory and emergent interactive narrative. International Journal of Continuing Engineering Education and Lifelong Learning, 14(6), p.506. doi:10.1504/ijceell.2004.006017.
Mazzone, M. and Elgammal, A. (2019a). Art, Creativity, and the Potential of Artificial Intelligence. Arts, 8(1), p.26. doi:10.3390/arts8010026.
Mazzone, M. and Elgammal, A. (2019b). Art, Creativity, and the Potential of Artificial Intelligence. Arts, 8(1). doi:10.3390/arts8010026.
Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K. and Galstyan, A. (2021). A Survey on Bias and Fairness in Machine Learning. ACM Computing Surveys, 54(6), pp.1–35. doi:10.1145/3457607.
Merleau-Ponty, M., Lefort, C. and Lingis, A. (1992). The visible and the invisible. Evanston, Ill. Northwestern University Press, p.96.
Metz, R. (2022). AI won an art contest, and artists are furious. [online] CNN. Available at: https://edition.cnn.com/2022/09/03/tech/ai-art-fair-winner-controversy/index.html.
Monserrate, S.G. (2022). The Staggering Ecological Impacts of Computation and the Cloud. [online] The MIT Press Reader. Available at: https://thereader.mitpress.mit.edu/the-staggering-ecological-impacts-of-computation-and-the-cloud/.
Oppenlaender, J. (2022). A taxonomy of prompt modifiers for text-to-image generation. [online] doi:10.48550/ARXIV.2204.13988.
Parisi, L. (2019). Critical Computation: Digital Automata and General Artificial Thinking. Theory, Culture & Society, 36(2), pp.89–121. doi:10.1177/0263276418818889.
Pearson, M. (2011). Generative art: a practical guide using processing. Shelter Island, Ny: Manning; London.
Plain, A., Eynon, R., Hjorth, I. and Osborne, M.A. (2022). AI and the Arts: How Machine Learning is Changing Artistic Work. Report from the Creative Algorithmic Intelligence Research Project. Oxford Internet Institute, University of Oxford, UK.
Poltronieri, F. (2022). Towards a Symbiotic Future: Art and Creative AI. The Language of Creative AI, pp.29–41. doi:10.1007/978-3-031-10960-7_2.
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C. and Chen, M. (2022). Hierarchical text-conditional image generation with CLIP latents. [online] doi:10.48550/ARXIV.2204.06125.
Rombach, R., Blattmann, A., Lorenz, D., Esser, P. and Björn Ommer (2021). High-resolution image synthesis with latent diffusion models.
Roose, K. (2022). An A.I.-Generated Picture Won an Art Prize. Artists Aren’t Happy. The New York Times. [online] 2 Sep. Available at: https://www.nytimes.com/2022/09/02/technology/ai-artificial-intelligence-artists.html.
Seele, P. and Schultz, M.D. (2022). From Greenwashing to Machinewashing: A Model and Future Directions Derived from Reasoning by Analogy. Journal of Business Ethics, [online] 178(4), pp.1063–1089. doi:10.1007/s10551022050549.
Suresh, H. and Guttag, J. (2021). Understanding Potential Sources of Harm throughout the Machine Learning Life Cycle. MIT Case Studies in Social and Ethical Responsibilities of Computing. doi:10.21428/2c646de5.c16a07bb.
van Eck, C. (n.d.). Gell’s theory of art as agency and living presence response. [online] Leiden University. Available at: https://www.universiteitleiden.nl/en/research/research-projects/humanities/gells-theory-of-art-as-agency-and-living-presence-response [Accessed 1 Nov. 2022].
- 1 LARPing as the algorithmic ouroboros.
- 2 This won’t follow the approach where I make clear what is and isn’t generative content as the text being input is my own and I see this as no different to using a grammar or spell-checker, or asking advice from a knowledgeable peer and qualifying their view.
- 3 https://beta.openai.com/playground
- 4 https://github.com/CompVis/stable-diffusion
- 5 https://openai.com/dall-e-2/
- 6 What is interesting with fragments [1] and [2] is that both impart high-level cognitive behaviours to machines — “understand”, “interact”, “respond” — as if the model were a mind rather than a statistical compression of human labour; the anthropomorphism is not innocent, since it is exactly what lets the human work inside the machine drop out of view. Something to bear in mind as we explore further.
- 7 https://deepdreamgenerator.com/
- 8 https://github.com/lucidrains/big-sleep
- 9 https://github.com/joel-simon/ganbreeder
- 10 https://openai.com/dall-e-2/
- 11 https://www.midjourney.com/
- 12 https://github.com/CompVis/stable-diffusion
- 13 https://runwayml.com/
- 14 If we review the Wikipedia history for the term (https://en.wikipedia.org/w/index.php?title=Prompt_engineering&action=history) we can see that the term doesn’t appear in Wikipedia until 2021 but in the research literature the first time we see the term used in a paper is 2018 (https://scholar.google.com/scholar?hl=en&as_sdt=0%2C5&q=prompt+engineering&btnG=).
- 15 https://lexica.art
- 16 https://promptbase.com
- 17 https://www.reddit.com/r/midjourney/comments/xqyn9e/prompt_engineering_is_an_art/
- 18 https://discord.gg/unstablediffusion; https://discord.gg/ymJMvgJY
- 19 https://www.reddit.com/r/StableDiffusion
- 20 Prompts are converted into high-dimensional vectors that are used to generate the images so if specific keywords are more weighted to one of the artists the images will skew towards their work that has been encoded in the model. Similarly, if an artist has more work encoded into the model then the weight of the keywords associated with their work will shift the image’s properties. These models are purely probabilistic and my explanation is an over-simplification but roughly allows us to understand the interaction between prompts and artist names.
- 21 This calculation was made using https://www.sust-it.net/energy-calculator.php and the highest kw/h value for my laptop’s energy consumption while running the processor and GPU at full performance. It ignores the original cost of training the model and the cumulative cost of its use on a global scale — the very costs the £3.40 is designed to make invisible.
- 22 https://drive.google.com/file/d/1w4YZB0tl40g7gw60nbnbT__ADxULIWCL/view?usp=sharing
- 24 These categorisations are to illustrate the broad differences in the artists generated by an LLM and my own qualitative impression of this process rather than as a true quantitative representation or study. The numbers are slightly off due to either dual-nationality, ambiguity or mixed-heritage and so artists being included twice in different categories — a slippage that is itself the point, since bodies rarely sit still inside the boxes these categories provide.
- 25 See the Wizard of Oz or any of the abrahamic texts.
- 26 We already have bread without a baker and candles without a candlestick maker.
- 27 I had a chat with a bot based on GPT-3 to explore this idea further and after initial disagreement, it finally agreed that artists used in prompts should receive a share of the profits made from their use.
- 28 The vagueness is not the practitioners’ failing so much as the infrastructure’s achievement: the “cloud” is engineered to be imperceptible, its whole rhetorical work being to make computation feel weightless and placeless, so that the substrate can be depended on without ever being felt.
- 29 https://makeavideo.studio
- 30 https://www.synthesia.io
- 31 Although, it isn’t like most national galleries are full of paintings of the wealthy in very similar styles.
- 32 Kudos (likes/follows), NFTs, payments for digital artefacts (prompts, models, plugins) or subscription payments to the individual.