What the laboratory studies
Longitudinal systems
What does a computational system do when it is allowed to keep going?
Most evaluation of computational systems is short and stateless: a task is posed, an answer is produced, the state is discarded. The laboratory studies the other regime — systems that persist for long periods inside a stable environment, where today's behavior is conditioned on everything that came before.
Persistence changes the questions. Stability, drift, and the slow emergence or decay of regularities are only visible over time, and only if the record is continuous enough to see them.
Falsification criterion. If a system's long-run behavior is fully explained by its short-run behavior, longitudinal study adds nothing. That would itself be a finding.
Developmental learning
Does a system that develops through experience differ from one that is configured?
A developmental system is one whose later capacities depend on its earlier history — on what it encountered, in what order, and what it did about it. The laboratory constructs worlds in which such histories can unfold, and asks whether anything develops that was not put there directly.
Whether intelligence, or any other complex behavior of interest, emerges from such systems is an open empirical question. The laboratory does not assume the answer, and does not claim to have found it.
Falsification criterion. If developmental histories produce no capacities beyond those attributable to initial configuration, the developmental hypothesis fails for that system and that world.
Memory and accumulated experience
When a system remembers, what is it actually retaining — and does it matter later?
Persistence is not the same as memory, and memory is not the same as learning. The laboratory is interested in distinguishing them: whether retained experience influences later behavior, whether that influence is appropriate to the situation, and whether it can be told apart from mere retrieval of the past.
Falsification criterion. If a system's behavior after a long history is indistinguishable from its behavior with that history erased, accumulated experience is not doing work.
Epistemic reliability
How does a system treat evidence, and can its treatment be trusted?
A system that acts on beliefs must acquire them, hold them, and sometimes give them up. The laboratory studies evidence handling and belief revision: whether a system's revisions track the evidence it was given, whether it is moved by evidence that should not move it, and whether it can be induced to hold a belief the evidence does not support.
This is also a question about the laboratory. An observer with a preferred outcome is a source of contamination, and the experimental design has to account for that.
Falsification criterion. If revision patterns are better predicted by the researcher's expectations than by the evidence presented, the reliability being measured is the researcher's, not the system's.
Reproducibility
Can an experimental history be replayed, and what has to be preserved for that to be possible?
Long-running experiments are hard to reproduce precisely because they are long-running: there is more history to preserve, more ways for it to diverge, and more temptation to summarize instead of retain. The laboratory treats reproducibility as an infrastructure problem first — what must be recorded, in what form, and with what guarantees — and only then as a methodological one.
Falsification criterion. Replaying a preserved history should produce the preserved record. Where it does not, the divergence is a defect to be reported, not a detail to be smoothed over.
Observation and intervention
How do you watch a system, and act on it, without becoming the thing you are measuring?
Observation is not free. Instruments consume resources, change timing, and can alter what they observe. Interventions — deliberate changes to the world or the subject — are sometimes necessary and always consequential. The laboratory records both, distinguishes them from each other, and treats the observer's effect on the experiment as a quantity to be measured rather than assumed away.
Controlled interventions are designed in advance, bounded, authorized, executed, and logged. An unrecorded intervention is a break in the experiment.
Falsification criterion. If the observed behavior differs materially between observed and unobserved conditions, the observation apparatus is part of the experiment and must be reported as such.
Evidence preservation and provenance
When a claim is made about what happened, can anyone check?
Experimental history is preserved as recorded, with provenance: what was captured, when, by which instrument, under which authority. Retained evidence can be identified by cryptographic hash so that a document published later can be tied to the record it describes. The record precedes the interpretation and outlives it.
Falsification criterion. A claim that cannot be traced to preserved evidence is, for the laboratory's purposes, not a result.
Experimental governance
Who is allowed to change the experiment, on what grounds, and how is that recorded?
Autonomous and semi-autonomous experimentation raises a governance problem: parts of the apparatus can propose or perform actions without a person in the loop. The laboratory's position is that autonomy must be paired with accountability — proposals are separated from authorization, authorization from execution, and every step leaves a record that a reviewer can audit afterwards.
Scientific integrity in this setting is a property of the process, not of the people. The process has to be one that cannot easily fool itself.
Falsification criterion. If an action affecting the experiment cannot be attributed to a recorded authorization, the governance model has failed for that action.
How claims are labelled
Whenever the laboratory discusses an experimental claim, it says which of these it is. A hypothesis does not become a fact because it would read better.
- KnownDirectly established facts about the laboratory, its software, its configuration, or an experiment.
- ObservedMeasurements or behaviors recorded during experiments.
- InferredInterpretations supported by observations but not directly measured.
- HypothesizedTestable explanations or predictions that have not yet been established.
- UnknownQuestions for which the available evidence does not justify a conclusion.