English Help Legal Sign Up Log In

Building a Markov chain bot to generate realistic email content

Most email outreach operations eventually hit the same wall: recipients recognise the same recycled templates within weeks, and reply rates crater as a result. Markov chain generation offers a way out by stitching together sentences that look fresh on the surface but still feel on-brand to the reader. The technique is older than most marketers realise, sitting inside the same family of probability models that power phone autocorrect and basic chatbots.

For Australian senders, the appeal is partly economic. Buying lists from brokers in Sydney or Melbourne keeps the cost per lead in AUD manageable, but the conversion gap between human-written copy and static templates makes automation worthwhile. A well-tuned chain can produce thousands of unique body variants from a single training corpus, letting a small team in Brisbane or Perth push volume without outsourcing to a copywriter.

Compliance still applies, and the Spam Act 2003 makes ignoring it expensive. The bot needs to respect unsubscribe language, sender identification, and consent records before anything goes out. Marketers who got fined by ACMA tend to be the ones who treated automation as a permission to skip the boring parts, so the technical build must include those guardrails from day one.

The chain itself does not care about grammar, intent, or persuasion — it only knows which word is likely to follow another. That limitation is actually useful for outreach, because a slightly awkward email often lands better in the inbox than a polished sales pitch. Readers in Parramatta offices or Adelaide coworking spaces are used to informal tone from real colleagues, and a Markov-generated message can sound exactly like that.

How Markov chains actually work for text

A Markov chain models a system as a series of states, where the next state depends only on the current one. Applied to text, each state is a word and the transitions count how often one word follows another in a training corpus. Bigrams use the previous one word, trigrams use two, and the order is the main dial that changes how readable the output becomes.

The simplest engine splits the corpus into tokens, counts every adjacent pair, and stores the result in a dictionary like {"thanks for": {"the": 4, "your": 2, "reaching": 1}}. To generate, the bot picks a seed, walks the dictionary, and at each step samples one of the following words weighted by frequency. Repeating that walk produces a string of words that mirrors the style of the source material.

Higher orders give smoother prose but explode the memory footprint. A trigram model trained on 500 marketing emails might weigh in at a few megabytes, while a five-gram model can hit 200 megabytes and still need a hash table to look up transitions quickly. The sweet spot for outreach usually lives at order two or three, where sentences still surprise but rarely collapse into nonsense.

Sourcing and cleaning a training corpus

The output is only as good as the input. Pulling a few thousand genuine B2B emails from a shared inbox gives the chain real vocabulary, and pairing that with newsletter copy produces more varied openers. A useful starting point is to scrape your own sent folder, then expand with public newsletters or industry reports to widen the vocabulary.

Cleaning matters more than volume. Strip email signatures, phone numbers, and reply chains because the chain will treat "Kind regards, John" as a complete sentence and start producing names mid-paragraph. Lowercase the whole corpus, collapse repeated whitespace, and keep punctuation attached to the preceding token so the chain learns sentence boundaries.

For Australians targeting local SMBs, mix in a few domain-specific samples — tradie quote follow-ups, real estate appraisal notes, accounting intake forms. The chain will pick up the casual register that Aussie recipients expect, including contractions and the occasional colloquialism that feels native to a Perth tradie or a Melbourne cafe owner.

A handy resource for pattern mining is corpersbook.com, which aggregates corporate contact structures and signing conventions across industries. Feeding its signature blocks through the cleaner gives the chain a feel for how different professions actually close emails, and that detail often makes the difference between a message that reads as a bot and one that reads as a tired salesperson.

Tuning entropy and choosing chain order

Entropy controls how predictable the output is. A chain with low entropy reproduces the training set almost verbatim, while high entropy produces nonsense. Outreach wants the middle: enough novelty to dodge spam filters and template detectors, but enough coherence that a recipient in a Surry Hills office will actually read past the first line.

Temperature is the usual knob. Divide the log probabilities by a temperature value before sampling; values above 1 increase randomness, values below 1 sharpen the distribution. Most operators settle on 0.8 to 1.2 for outreach, with rare excursions above 1.5 for A/B tests on cold segments.

Backoff handles unseen n-grams. When the chain hits a trigram it never saw, fall back to a bigram of the last two words, and finally to a unigram. Without backoff, a single unseen word kills the entire sentence, and the bot starts emitting empty strings that the sender tool rejects.

Order Memory per 10k emails Readability Spam-filter evasion Best use case
1 (unigram) ~2 MB Poor Low Subject lines only
2 (bigram) ~15 MB Decent Medium Short follow-ups
3 (trigram) ~80 MB Strong High Cold outreach bodies
4 (quadgram) ~400 MB Very strong Lower Risk: triggers duplication detectors
5+ 1 GB+ Variable Unstable Not recommended

Logging is mandatory. Save the seed, the parameters, and the first 50 outputs of every batch so you can replay a campaign that converts well. Australian senders in particular benefit because time zones in AEST make iterative testing slow — a half-finished chain rerun at midnight on a Friday is not the same experiment as a Monday morning batch.

Writing the minimal generator in Python

The whole engine fits in under a hundred lines of Python. Read the file, build the transition dict, pick a seed, and walk. The collections.defaultdict structure handles missing pairs gracefully, and random.choices with weights handles the sampling. No external libraries are required for the core, though numpy speeds up heavy corpora.

Pseudocode looks like: tokenise the text, slide a window of size n across the tokens, append the next word to the dict under the current window as key, then store as values. For generation, pick a random seed from the dict keys, emit it, then for each step look up the current state, sample a follower, emit it, and shift the window.

One trap is infinite loops on short states. If the chain lands on a key whose only follower is itself, the output stalls. Guard against it by capping sentence length, forcing a period every 30 words, or seeding the walk from a list of common sentence starters. Adding a small chance of jumping to a new random key keeps the prose moving when the corpus is thin.

Wiring the generator into an outreach pipeline

The output of the chain is plain text, so any sender that accepts a body field works. Pipe the generator into an SMTP relay, an SMTP service like Mailgun, or a private SMTP that supports high volume from Australian data centres. Telstra-grade IP reputation matters more than the cleverness of the prose, so warm up new sending domains over two weeks before the bot's output goes anywhere near them.

Personalisation layers on top. Run the chain per recipient by injecting first name, company, or city into the seed prompt, then generate. A seed like "Hi Sarah, just following up on the" produces a paragraph that mentions Sydney or Brisbane without the chain ever seeing that data, which keeps the training corpus generic and reusable across lists.

Throttle the sends to mimic Australian business hours. AEST runs from UTC+9, with daylight saving pushing parts of the country to UTC+11. Schedule the bot to send between 9 am and 4 pm local, with heavier volume mid-morning because Australian open rates peak then. Sending at 3 am local time in Melbourne is the fastest way to land in junk.

Track replies, not opens. Open tracking is unreliable after Apple's MPP rollout, but reply content tells you whether the chain is producing language that recipients accept as human. Save every reply, feed the good ones back into the corpus, and retrain monthly. The bot improves faster this way than any parameter tuning alone.

Common pitfalls when training a Markov chain for email

  • Feeding HTML bodies without stripping tags
  • Including reply chains that introduce multiple speakers
  • Treating names and email addresses as part of the vocabulary
  • Ignoring punctuation, which collapses sentence detection
  • Training on a corpus under a few hundred unique sentences
  • Skipping deduplication before counting transitions

Australian compliance checks before scaling the bot

  • Confirm every recipient has opt-in evidence stored for ACMA inspection
  • Include a working unsubscribe header in every generated body
  • Identify the sender with a real physical address, not a P.O. box only
  • Honour unsubscribe requests within five business days
  • Keep records of consent for at least 24 months
  • Avoid subject lines that misrepresent the body content

A Markov chain is not a magic wand, but for senders who already have clean lists and good infrastructure, it quietly removes the bottleneck that template fatigue creates. Engineers in Adelaide and Perth tend to underrate the value of a clean seed list because their local markets are smaller, though a tight chain trained on the right corpus can outperform a sloppy campaign ten times the size, and reply data from real Aussie prospects keeps the model honest in ways that scraped international samples never can.