Research
Hypotheses published before the results.
Design-research from inside thirdwurld. Each entry states what we think will happen, how it is being measured, and what result would prove it wrong.
The demo shows what already runs. This is the other half: the claims that are not settled yet. Everything here is written to be falsifiable, so a null result is a publishable outcome rather than something to bury. Where a document is a protocol rather than a finding, it says so at the top. Terms such as preference, memory, and choice describe system behavior, not subjective experience.
-
PDF
A resident that stays still can look like a model limitation, but locomotion is a pipeline: readiness, human occupancy, scheduling, path construction, collision validation, movement input, and arrival. The study refuses to treat a visible correlation as a causal answer before those layers are separated.
The paper tests three distinct questions: whether models can produce a valid bounded movement action when instructed, whether they select walking when quiet time and other actions are available, and whether the authoritative execution layer completes an accepted action. A deterministic no-model scheduler is included as a control.
The central prediction
In the existing runtime, ordinary movement is server-scheduled, so model tier should not explain walking after readiness, occupancy, route availability, and selected activity are controlled. If a model difference emerges after movement is exposed as a bounded action, the study identifies whether it is in action selection, recovery, or physical execution.
Pre-committed rules
Decision Requirement Capability advantage ≥ 10 points, CI excludes zero Higher-tier default ≤ 3x cost per completed action Production execution none during this study Ordinary movement restriction never based on casual observation Boundary
The protocol measures outputs, server outcomes, and visible movement. "Voluntary" means the model selected an allowed action under a specified state.
Keywords
embodied AI action grounding tool use agent reliability virtual worlds pre-registrationNo trial, provider call, resident behavior change, or production deployment has occurred. The protocol exists so a null or inconvenient result is still publishable.
-
PDF
Large language models can give characters fluent speech, but fluency is not evidence that anything happened. A resident that says it found a coin on the north walk is producing text, not testimony. The paper argues the measurable property of a persistent world is not believability but verifiability: whether a resident's account of its own past can be checked against authoritative world state, and whether verified history changes what it does next.
Believability is rejected as an engineering target because it is confounded by novelty and anthropomorphic projection, it cannot be measured without humans, and it can be maximised by a system that lies fluently. A resident that confabulates a rich history will outscore one that accurately reports a thin one.
The intervention stays deliberately small. The server creates one scarce object, the Wayfarer's Coin. A resident may choose exploration over routine, traverse a collision-grounded route, and claim the object through an atomic transaction that writes an
object_discoveredevent to an append-only ledger. Later behaviour may retrieve it.The two load-bearing hypotheses
- H1, grounded coherence. Requiring a valid route, arrival and atomic claim before a discovery can be narrated raises the proportion of resident narrations fully supported by the ledger.
- H2, durable continuity. A discovery written to durable memory produces later references that are more accurate against the ledger, and more contextually appropriate, than an identical route and claim with no memory write.
Both are testable with zero human participants. A third hypothesis on human return is marked exploratory, because a power calculation shows current session volume cannot detect a plausible effect. The paper states that limit rather than reporting an underpowered result as a finding.
Some of the pre-committed thresholds
Criterion To proceed Ledger-support rate, grounded arm ≥ 0.95 Ledger-support advantage over ungrounded ≥ 0.20 Retrieval accuracy, memory over no-memory ≥ 0.15 Claims without complete route evidence zero p95 end-to-end claim latency < 2.0s Boundary
Resident narration is evaluated as a claim against world records, not as evidence of experience.
Keywords
generative agents agent memory grounding verifiability embodied AI persistent virtual worlds hallucination measurement human-computer interactionPre-deployment protocol. No results are reported and no pilot has been run. Every numeric threshold above is committed in advance of data collection, which is the point of publishing it now rather than after.
-
PDF
Multi-agent systems are routinely reported to show social transmission: agents tell each other things and information spreads. Whether what spreads is true is rarely measured in persistent-world agent systems with machine-checkable ground truth. This paper applies the transmission chain method from the study of human cultural transmission to AI residents.
One world event is verified by the server ledger. The resident who witnessed it tells a second, who tells a third, along a controlled chain. Because the origin is machine-readable, every downstream account can be scored automatically against ground truth. The headline measure is the event half-life: the number of hops at which fewer than half of checkable assertions still hold up.
Four directional predictions
- H1, fidelity decay. Ledger support falls monotonically with hop distance, by at least 0.25 absolute from hop 0 to hop 3.
- H2, differential field decay. Fields fail in a set order. Claimant identity goes first, exact position next, coarse region survives longest.
- H3, attribution drift. Without provenance tagging, residents progressively misattribute the event to themselves, producing confident first-person accounts of things they did not do. Predicted above 0.15 by hop 3.
- H4, the fix. Tagging every memory with whether it was witnessed or heard, and from whom, cuts misattribution by at least half.
The second clause of H4 prevents a misleading improvement: a system can avoid false attribution by becoming vague. The protocol tracks the unverifiable share alongside accuracy.
Some of the pre-committed thresholds
Criterion Threshold Hop 0 ledger-support, all arms ≥ 0.95 Decay, hop 0 to hop 3 ≥ 0.25 absolute Attribution drift, untagged arm at hop 3 > 0.15 Provenance effect on drift ≥ 50% reduction Vagueness guard, unverifiable share ≤ +0.10 The uncomfortable possibility
The paper states plainly that the result may be that transmission between residents is too lossy to carry shared history at all, which would remove a feature from the roadmap. It argues that is still the most valuable outcome available, since the alternative is shipping social propagation and diagnosing the same problem later through incoherent world history.
Boundary
The study scores accounts against a ledger; it does not infer belief or understanding.
Keywords
multi-agent systems transmission chain agent memory provenance misinformation grounding generative agents persistent virtual worldsPre-deployment protocol. No results are reported. Builds directly on the verifiable memory protocol above, which establishes the ledger this study scores against.
What belongs here
Protocols before they are run, results once they are in, and occasional notes where a design decision needed an argument rather than an opinion. Claims that cannot be tested belong on the demo, not on this page.