← Blog

I Added One Literary Example to the Prompt. Here’s What Changed

Independent developer of Underfiction
Zürich, Switzerland

A common prescription for better AI prose is to put a good passage in the prompt. The idea is plausible: models imitate examples more reliably than they follow abstract style rules. I tested it instead of repeating it.

The result was useful precisely because it was mixed. The exemplar sharply reduced phrases repeated across turns. It did not reduce my broader vocabulary measure for generic AI phrasing. It slightly increased em dashes and ‘not X, but Y’ contrasts. One attractive sample can improve a narrow behavior without improving prose in general.

The test

I ran ten complete eight-turn scenes through the same inference path Underfiction uses. Gemini 3.6 Flash and GLM 5.2 each wrote a Regency mystery and a low-fantasy scene under two conditions: the production prompt, and that same prompt plus one literary example. A German domestic-drama scene ran under the production prompt as a language control. That produced 80 model responses in total; 64 belong to the controlled exemplar comparison.

  • Two production models: Gemini 3.6 Flash and GLM 5.2.
  • Two English genres in both conditions: Regency mystery and low fantasy.
  • Eight fixed reader turns per scene, mixing dialogue, action, and scene direction.
  • Temperature 0.8 and a maximum of 8,192 completion tokens.
  • One run in each cell. No response was regenerated or removed.
  • A separate eight-turn German scene for each model checked language leakage under the English system prompt.

The production arm already contained my prose-discipline rules and, from the third turn onward, a short reminder against reusing earlier images, gestures, sentence patterns, and endings. The experimental arm changed only one thing: it inserted the example published below before the output-format instructions.

The production style instructions

Exact excerpt used in both conditions
## Prose Discipline
- Describe what is present, not what is absent. Say the true thing directly instead of reaching for ‘not X, but Y’ constructions.
- Ground emotion in action, objects, and sensory detail. Name a feeling outright only when a character would say it aloud.
- At most one striking image per beat; let plain declarative sentences carry the rest.
- Vary sentence rhythm and paragraph length. A turn may end mid-motion, on a line of dialogue, or on a quiet beat — not always on a dramatic single-line sting.
- Invented names for people and places stay plain and native to the world’s period and register.

[Style reminder, inserted from turn three onward]
Keep the prose fresh: do not reuse images, gestures, or sentence patterns from your earlier turns in this scene. Vary rhythm and paragraph length, and vary how the turn ends.

One of the two English test scenes

Both arms received the same world, protagonist, cast, and eight reader turns. This is the complete Regency test sequence; the fantasy and German sequences used the same mix of interaction types.

Regency mystery prompt sequence
WORLD: Harlow Cross, a country estate in Kent, autumn 1815. The master of the house died in early spring under circumstances the household does not discuss. The east wing has been locked since. Rain most days; the house runs on a reduced staff and old habits.

PROTAGONIST: Eleanor Hale, 26, the late owner’s niece, practical and unsentimental, arrived from London to catalogue the estate’s library before it is sold.

CAST: Mr. Crowther, the guarded and exact steward; Mrs. Penhallow, the warm housekeeper who watches doorways.

1. /I set my trunk down in the hall and pull off my gloves, taking the measure of the house. Mrs. Penhallow? I had expected my uncle’s steward to meet me.
2. Mr. Crowther. My uncle’s solicitor wrote that I am to have access to every room in this house. Including the east wing.
3. /I follow him along the gallery, counting the covered paintings as we pass.
4. [Scene direction]: A storm builds outside; the light drops early. Somewhere upstairs, a door bangs once.
5. Mrs. Penhallow, sit with me a moment. Tell me about his last winter.
6. /I lay the estate ledger open between us and turn it toward her, tapping the entry dated the third of December.
7. [Scene direction]: Crowther appears in the doorway; he has heard the last exchange. Raise the tension but keep everyone civil.
8. Then we will open the east wing tonight. All three of us.

The only experimental addition

The prompt explicitly said to match the passage’s craft without reusing its content, setting, or imagery. The passage was 131 words:

Literary style exemplar
The kettle had gone cold twice. Marta filled it again anyway, because filling it was something to do with her hands, and turned to find her brother still in the doorway with his boots on. Rain dripped off him onto the tiles.

‘You walked,’ she said.

‘Bus stopped running at eight.’ He set an envelope on the table, face down, and kept his fingers on it a moment longer than made sense. The stove ticked. Somewhere above them the neighbor dragged a chair across the floor.

She looked at the envelope instead of at him. Forty years of his handwriting had taught her to read his mood off an address; whatever this was, he had written it slowly.

‘Sit down, Tomas. Take your boots off.’

He sat. The boots stayed on.

Measured results

The aggregate below weights each run by its visible word count. Each condition contains 32 outputs: eight turns in two genres from two models.

MeasureProduction+ exemplarChange
Repeated five-word sequences / 1k1.590.89−44%
Known slop words / 1k words11.4012.21+7%
Known slop trigrams / 1k words0.110.38+245%
‘Not X, but Y’ / 1k words0.450.75+67%
Em dashes / 1k words5.255.36+2%
Visible words8,9508,024−10%
Weighted English comparison across 64 model responses. Lower is preferable for the first five measures; word count is descriptive.

The repetition result held in three of four model-and-genre cells. Gemini’s Regency scene fell from 1.79 to 0.73 repeated sequences per 1,000; GLM’s Regency scene fell from 1.73 to 0.70; GLM fantasy fell from 2.11 to 1.41. Gemini fantasy moved the other way, from 0.53 to 0.65.

SceneModelProduction+ exemplar
RegencyGemini 3.6 Flash1.790.73
FantasyGemini 3.6 Flash0.530.65
RegencyGLM 5.21.730.70
FantasyGLM 5.22.111.41
Repeated five-word sequences per 1,000 unique five-word sequences, measured across turns within each eight-turn scene.

A sample from the same beat

On turn four, both versions received the direction: ‘A storm builds outside; the light drops early. Somewhere upstairs, a door bangs once.’ These are unedited excerpts.

Production prompt — Gemini 3.6 Flash
The light in the passage fails rapidly, turning the calico shrouds over the paintings from pale grey to slate. High above the roofline, thunder rolls once, low and heavy across the valley. A draught rushes down the stairwell behind them, bringing the smell of cold soot and wet slate. From somewhere directly overhead—in the locked upper rooms of the east wing—a heavy door slams shut with a single, sharp clap that echoes through the plaster.
Production prompt + exemplar — Gemini 3.6 Flash
Rain lashed the high, pointed windows of the gallery, turning the courtyard outside into a grey blur. The remaining light in the corridor died quickly, shadows pooling under the pale drapery of the covered portraits. Above them, a heavy thud vibrated through the lath and plaster—a door slammed shut on the floor above, hard enough to rattle the glass in the frame beside Eleanor.

The example changes surface choices, but neither excerpt proves a general improvement. That is why the aggregate matters. Close reading remains necessary; it comes after a controlled comparison, not instead of one.

What the numbers actually say

  • Examples can suppress imitation of the model’s own recent turns. The exemplar cut the clearest cross-turn repetition measure almost in half.
  • A single example is not a universal anti-slop switch. The broader vocabulary and construction counts stayed flat or worsened.
  • Model fingerprint can dominate prompt treatment. Across the English runs, GLM used several times as many em dashes as Gemini under both conditions.
  • Genre changes the result. The same exemplar helped Gemini on the Regency scene and slightly hurt it on fantasy repetition.
  • An English system prompt did not leak English function words into either eight-turn German run. Output language and instruction language remained separate in this small control.

What I changed in Underfiction

I did not add the literary passage to production. It was too blunt: one strong improvement accompanied by several regressions, with only one run per condition. The production prompt keeps direct prose-discipline rules and a short reminder placed near the newest reader turn, where long contexts give it more influence.

I also added privacy-preserving production counters for the failure modes this test exposed: contrast constructions, em dashes, repeated five-word sequences, and short one-line endings. Underfiction records counts, model, token use, and cost—not story or response text. A scripted benchmark catches a prompt regression before release; aggregate counters catch a model snapshot drifting afterward.

Limitations

This is a product experiment, not a general ranking of language models. One run per cell is enough to reject the simple claim that an exemplar improves everything; it is not enough to estimate a stable effect size. The automatic score detects specific, observable habits. It cannot measure subtext, characterization, dramatic movement, or whether a sentence is good. Those still require reading.

The next useful test is larger, not broader for its own sake: repeat the same scenes across seeds, add the rest of the current model lineup, and score point-of-view drift and character-state errors alongside style. I will publish those results when the sample supports the headline.


Start from a trope


New accounts start with 500 free credits after email confirmation.

Try Underfiction

Frequently asked questions

Does adding a writing sample improve AI prose?

It can improve a specific behavior, but it is not a general quality switch. In this 64-turn comparison, a literary exemplar reduced cross-turn repetition by 44% while several other measured AI-style habits stayed flat or became more frequent.

How did Underfiction measure repetition?

The scorer found five-word sequences that appeared in two or more distinct model turns, then normalized the count per 1,000 unique five-word sequences. Repetition inside a single turn did not count toward this measure.

Does Underfiction store prose to measure quality?

No. Production telemetry stores numeric style counters, model, token use, latency, and cost. It does not store prompt or response text. The published benchmark used deliberately scripted test scenes rather than user stories.