07 — Falsifiability
Make PULSE wrong-able.
The architecture is not the bottleneck. Phase 2 is the contract an engineer can try to break: formal objects, causal timing, calibrated uncertainty, synthetic adversaries, and one experiment that moves ΔL.
What is locked
Not a redesign. The same three layers, now with objects that can fail a test.
| Object | Definition |
|---|---|
| zᵤ | Latent class in {human, human+agent, declared-agent, service, coordinated}. |
| pᵤ(t) | P(zᵤ | Hᵤ(t)). A posterior, not a badge. |
| Lᵤ | A functional of pᵤ — mass on independently acting human, discounted mixed. |
| (μ_L, σ_L) | Point and width. Cold start is a wide interval, not a low score. |
| q_t | Causal multiplier. Only H_<t is legal. |
| CIᵤ | Predictive lift of the cluster residual after the stimulus shown. |
Posterior, not a binary
Do not write Lᵤ = E[independent human | Hᵤ]. That collapses the taxonomy the product already needs. Declared agents should be allowed to participate without faking a pulse.
Causal timing
The reaction field of event E develops after people begin reacting. A ranker that consumes P(E) in the same instant it scores the first viewers is using the future. Streaming contract:
Early seconds of a post have almost no event field. Author-side L and edge history carry the score. Event EKG updates distribution as the wave develops, and feeds offline ΔL. It does not give the first ranker omniscience.
Compatible with Phoenix candidate isolation: q_t is an author/edge/lagged-event feature. It does not require candidates to attend to each other.
Counterfactual independence
Alignment is not proof. Newsrooms, stan accounts, and earthquake threads couple on a public stimulus. The interesting object is whether the cluster residual still predicts account u after you know what u actually saw.
Mutual information is the target. The estimator at ranking scale is predictive lift, not a closed-form I(·;·|·):
High CI is evidence of a shared unobserved controller or objective. It is then typed with the rest of the taxonomy. CI is not the new bot score. Journalists living in a breaking-news cluster will light it up; they are not a farm until the rest of the posterior says so.
Calibration
Internally — never publicly — every liveness estimate carries a width:
A point of 0.72 with that interval is “do nothing.” Shrinkage toward a weakly informative prior on cold start is the product rule. If you cannot say how sure you are, you do not multiply the score.
The experiment worth running
- Event-level EKG on a slice of high-reach posts.
- Discount swarm-shaped engagement with causal q_t. Do not hang the author when the author is independently alive and the audience is not.
- Measure ΔL, conversation depth, response diversity, residual coupling, and CI lift against a holdout. Not raw click volume.
- Ablate: jitter, persona, clone, independent swarm, human farm. False-positive sets: night shift, journalist, scheduled posts, accessibility, non-native.