Skip to content

06 — Adversaries

Cost of imitation rises. It is not infinite.

Once a metric exists, generators train against it. Design for the swarm that looks independently human at the account layer and coupled at the objective layer — independent agents, independent clocks, a shared objective.

AttackWhat it doesWhere it should die
Jitter botUniform random delays, maybe typos.Layer 1. T is unconditional.
Persona botLLM + sleep schedule + slang.D and G fail to cohere across weeks.
Behavioral cloneImitate one scraped human across timescales.Layers 2–3 when many clones share residual coupling.
Independent agent swarmSeparate models, separate clocks, shared objective.CI lift and ΔL. This is the real enemy.
Human farmPaid people clicking. Pulse looks alive.Economic regularity, device reuse, payment. Not heartbeat alone.

What “train against it” looks like

If “irregular delays + typo + late reply” becomes the public template, the next generation samples that distribution. Account-history features survive text rewriting better than content features. They will not survive a well-funded imitation of your heartbeat forever. That is why the formula stays private, the output stays a prior, and the privileged instrument is the event field — a population reaction the attacker cannot fully observe from one account.

X already withholds some botmaker rules from the public algorithm repo for this reason. A heartbeat implementation should do the same: publish the philosophy and the consumption contract, not the exact coupling estimator.

Counterfactual independence as the hard test

A swarm that has independent clocks and original sentences will pass unconditional jitter checks. It still has a shared unobserved controller. After you condition on the stimulus each account actually saw, knowing the rest of the cluster still predicts the next act. That predictive lift is CI. High CI is not guilt — newsrooms will light it up — but it is the quantity the independent-agent swarm cannot cheaply drive to zero.

False positives are part of the threat model

A system that cannot tell a night-shift nurse from a farm is not ready to touch ranking. Paul Graham’s replies being hidden as possible spam is the cartoon of this error: polished language is not a bot. Power users, people with phones glued to them, journalists living in the feed, and accessibility tools all distort T, D, and E. Wide σ_L is how those accounts survive first contact with q_t.

The lab lets you fire a viral event in a swarm world and compare the pulse to an organic square. It will not substitute for a holdout of live campaigns. The experiment list is in Falsify.