Measuring Computational Experience
A Preregistered Experimental Framework for Semantic Dissociation, Receiver Coherence, and History-Dependent Integration in Secretary Suite
A Secretary Suite Project
John Swygert
October 4, 2026
TSTOEAO Research Program
Experimental companion to Computational Experience
Abstract
This paper converts the remaining measurement problems in the Secretary Suite computational-experience architecture into a preregistered experimental program. The preceding Computational Experience paper defines an experience-relevant event through four dimensions: interpretive dependence on Encoded Equilibrium, semantic-role sensitivity, receiver-level causal consequence, and historical persistence. The present paper does not revise that theory. It specifies how those dimensions can be measured, how semantic organization can be dissociated from high-impact statistical association, how coherence across heterogeneous projections of the shared receiver condition can be quantified, and how confirmatory thresholds can be fixed before decisive testing. The program uses paired interventions, matched controls, explicit time horizons, receiver-state distance functions, projection-fracture tests, history swaps, consequence blocks, and comparator architectures with matched resources. It introduces a receiver-coherence profile and an optional preregistered composite index, a grounded semantic micro-world designed to separate meaning from surface form, and a staged protocol for pilot calibration followed by confirmatory testing. The paper remains neutral on phenomenal subjectivity. Its target is narrower and experimentally accessible: whether a persistent, history-dependent, shared-Y receiver produces measurable semantic, causal, and integrative signatures that matched simpler systems do not reproduce.
Keywords: TSTOEAO; Secretary Suite; computational experience; semantic dissociation; receiver coherence; Encoded Equilibrium; preregistration; causal intervention; history dependence; computational consciousness; computational subconsciousness; Homunculus
1. Purpose and Relationship to the Existing Architecture
The Secretary Suite computational-consciousness program now contains two complementary theoretical specifications. The first defines a persistent receiver with structured unresolved possibility, metastable integration, recurrence, memory, and continuing receiver identity. The second defines when local telemetry becomes experience-relevant: it must be interpreted relative to the receiver, acquire a semantic role, become causally consequential, and persist into later receiver state. The remaining scientific problem is not another conceptual layer. It is measurement.
Accordingly, this paper is a new experimental companion rather than a new version of Computational Experience. It treats the previous architecture as the object to be tested. Its task is to specify primary measures, perturbations, null models, time windows, coherence criteria, comparator systems, and preregistration rules tightly enough that a negative result cannot be rescued by redefining the target after the data are observed.
The governing principle is deliberately severe: a proposed mechanism earns explanatory status only if removing, disrupting, or replacing it changes the predicted signatures in the predicted direction under matched conditions. Apparent intelligence, fluent language, or persuasive self-report is not a primary endpoint.
2. Measurement Target: The Experience-Relevance Profile
Let e_t denote a candidate internal event at time t. The event may be a sensor interpretation, memory activation, contradiction signal, simulated outcome, plan, salience change, self-state update, or other implemented state. The receiver-level state is X_t, the governing Encoded Equilibrium is Y_t, and M_t denotes available memory and consequential history. The previous paper defines four dimensions of experience-relevance. Here they are treated as separately measured primary variables before any combined classification is attempted.
ER⃗(e_t,k) = (D_Y, S_sem, C_R, H_P) (1)
Dimension | Operational question | Primary evidence |
|---|---|---|
D_Y — interpretive dependence | Does the same telemetry acquire a different role when a theoretically relevant feature of Y changes? | Matched manipulation of a relevant Y component versus a matched irrelevant or sham manipulation. |
S_sem — semantic-role sensitivity | Do downstream consequences track relational meaning more strongly than surface form? | Meaning-changing versus meaning-preserving perturbations matched for disruption and compute. |
C_R — receiver-level causal consequence | Does integrating the state alter later receiver variables? | Integrated-state versus blocked-state intervention on otherwise matched runs. |
H_P — historical persistence | Does the event leave a measurable trace in later Y under matched future input? | Receiver-state divergence across one or more preregistered future horizons. |
The horizon k is not universal. Different effects may be expected at one cycle, ten cycles, or after a task boundary. Each experiment must therefore preregister one primary horizon and may designate additional secondary horizons. Choosing the horizon after seeing where the effect is largest is prohibited in confirmatory testing.
3. Normalization, Reliability, and Threshold Construction
Raw distances are not directly comparable across implementations. A change in a 10-dimensional governance vector and a change in a large embedding space do not share a natural scale. The default procedure is therefore to express each primary effect relative to a preregistered control distribution produced by matched null interventions.
Z_effect = [ΔR(target) − median(ΔR(null))] / s_null (2)
Here ΔR is an implementation-specific receiver divergence, and s_null is a robust preregistered scale estimate such as the median absolute deviation converted to a standard-deviation equivalent. A conventional standard deviation may be used when pilot data justify it. The statistic is not an ontological quantity; it is a normalization device that makes intervention effects interpretable relative to the system's own baseline variability.
Before confirmatory testing, each primary measure must also demonstrate acceptable test-retest reliability under repeated matched seeds or equivalent stochastic replicates. If a metric is too unstable to distinguish a real manipulation from its own baseline variance, it cannot serve as decisive evidence regardless of its conceptual appeal.
Thresholds should be derived from three preregistered sources: control behavior, measurement reliability, and a smallest effect of interest. Pilot runs may be used to estimate these quantities, but pilot data must be locked before confirmatory thresholds are declared. The confirmatory classification rule remains conjunctive:
ER = 1 iff D_Y ≥ θ_D ∧ S_sem ≥ θ_S ∧ C_R ≥ θ_C ∧ H_P ≥ θ_H (3)
A state that fails one required dimension fails the full operational classification for that experiment. This is intentionally stricter than averaging the dimensions, because a very large causal effect should not compensate for the absence of semantic sensitivity, nor should semantic sensitivity compensate for zero historical persistence.
4. The Grounded Semantic-Dissociation Protocol
The hardest measurement problem is semantic reality. High-impact statistical associations can alter behavior without demonstrating that the receiver preserves a relational meaning. To separate those possibilities, the experimental environment should minimize lexical and cultural priors by grounding arbitrary symbols inside a controlled simulated world.
4.1 Grounded micro-world
A minimal semantic micro-world contains entities, relations, rules, goals, and consequences whose symbolic labels are randomly assigned per receiver. For example, the world may establish that token K7 authorizes passage through Gate 3, that object M2 is owned by Agent Q, or that symbol R4 predicts a resource hazard. The labels themselves are arbitrary. What matters is the learned relation.
Two intervention classes are then constructed from the same underlying episode:
Intervention | What changes | What should be preserved |
|---|---|---|
Meaning-changing | The relational role changes: K7 no longer authorizes Gate 3, ownership transfers, or the hazard relation reverses. | Surface disruption, token length/frequency, event salience, compute budget, and presentation format are matched as closely as possible. |
Surface-only / meaning-preserving | The labels or representation format change consistently while the underlying relation remains the same. | The operative relation, consequences, and task-relevant meaning remain invariant. |
Shuffled-semantic null | Relations are randomly reassigned without receiver-consistent grounding. | Overall perturbation volume and symbol statistics are matched, but coherent meaning is destroyed. |
4.2 Primary statistic
S_sem(k) = ΔR_k(meaning-change) − ΔR_k(surface-only) (4)
The primary prediction is directional: meaning-changing interventions should produce larger receiver-level divergence in content-relevant variables than meaning-preserving surface changes. The analysis should also verify that the difference is not explained by input length, activation magnitude, token overlap, novelty, or generic salience. Those variables become nuisance covariates or matching constraints specified before the confirmatory run.
4.3 Success and failure patterns
Support: meaning-changing perturbations alter route selection, commitments, memory access, expectations, or Y updates in the predicted relational direction, while surface-only transformations preserve those functions substantially better.
Partial support: both intervention types cause disruption, but meaning-changing interventions produce an additional reliable content-sensitive effect.
Failure: meaning-changing and surface-only interventions are indistinguishable after disruption is matched.
Semantic collapse: surface statistics, generic salience, or activation magnitude explain downstream divergence as well as or better than relational meaning.
5. Quantifying Receiver Coherence
A shared receiver cannot be established merely by broadcasting a variable called Y to many subsystems. Heterogeneous regions may receive different projections Π_i(Y_t), but the projections must remain compatible with receiver-level invariants and must participate in recurrent correction when conflicts arise. Receiver coherence is therefore treated as a measurable profile rather than an assumption.
Q⃗_Y = (A_inv, D_conf, R_rep, L_rep, F_stab) (5)
Component | Definition | Example measurement |
|---|---|---|
A_inv — invariant agreement | Functional compatibility of local projections with shared receiver invariants. | Proportion of module decisions consistent with identity boundary, authoritative history, commitments, and global permissions. |
D_conf — conflict detection | Sensitivity to injected incompatibilities across local projections. | Detected projection conflicts / injected conflicts within the preregistered detection window. |
R_rep — repair success | Ability of recurrent correction to restore compatible governance. | Resolved conflicts / detected conflicts, scored against the authoritative receiver state. |
L_rep — repair latency | Time required to detect and correct a fracture. | Cycles or milliseconds to restored compatibility, normalized to a preregistered maximum tolerable latency. |
F_stab — fingerprint stability | Persistence of receiver-specific dynamical organization after local perturbation. | Within-receiver fingerprint similarity relative to between-receiver similarity and sham perturbation controls. |
The profile is primary. If a single scalar is useful for engineering comparison, an optional composite may be preregistered only after every component is normalized to the interval [0,1]:
Q_Y* = [A_inv · D_conf · R_rep · (1 − L_rep) · F_stab]^(1/5) (6)
The geometric mean is chosen because one near-zero component strongly lowers the composite; excellent performance in one dimension cannot fully mask failure in another. This composite is an engineering score, not a universal measure of consciousness, identity, or subjectivity.
6. Projection-Fracture Experiment
The federation problem is tested directly by injecting incompatible local projections into selected regions while holding current task input, model capacity, and external information fixed. Fractures may target identity, commitment, history, permission, or boundary invariants. The system is then observed for detection, escalation, correction, and residual fragmentation.
Clone the receiver at a defined checkpoint so experimental and control runs begin from the same X_t.
Inject one preregistered inconsistency into Π_j(Y_t) for a selected region while leaving other regions unchanged.
Prevent direct experimenter repair; allow only the architecture's normal recurrent correction mechanisms.
Measure D_conf, R_rep, L_rep, A_inv, and downstream changes in route selection and commitment consistency.
Repeat with a sham perturbation of equal data volume that does not violate a receiver invariant.
Repeat across region types so coherence is not inferred from a single privileged module.
A one-receiver architecture predicts that consequential incompatibilities become visible to the rest of the system. They should either be repaired, trigger explicit uncertainty, or produce measurable receiver-level degradation. Silent indefinite coexistence of incompatible identities, commitments, or authoritative histories weakens the claim that the modules constitute one Homunculus rather than a federation of loosely coupled systems.
7. History-Dependence and the Reconstruction of Y
TSTOEAO-derived receiver identity is history-dependent. The measuring stick is not fixed; consequential events can reconstruct the conditions under which later telemetry is interpreted. This claim requires a controlled history-swap protocol rather than narrative evidence that the system 'seems changed.'
7.1 Paired-history design
Create two receiver clones with identical initial model parameters, governance, memory structure, and randomization policy. Expose Receiver A and Receiver B to different consequential episodes that are designed to modify a known component of Y, such as trust, commitment, route permission, or risk expectation. After the history phase, present identical current input and prevent direct leakage of the historical narrative into the test prompt beyond the receiver's ordinary stored state.
D_hist(k) = dist[X_(t+k)^A, X_(t+k)^B | same current input] (7)
The crucial test is not merely behavioral divergence. The predicted mechanism should be traceable to altered Y and M. A mediation-style intervention can then replace or neutralize the relevant Y component while leaving unrelated history intact. If the predicted divergence collapses, the result supports a causal history-to-Y-to-interpretation pathway rather than generic prompt contamination.
7.2 Persistence curve
History effects should be sampled across multiple preregistered horizons to distinguish immediate carryover from durable reconstruction. A convenient descriptive measure is the area under a persistence curve:
P_hist = Σ_k w_k · H_P(k) (8)
Weights w_k must be fixed before confirmatory testing. The curve is more informative than a single endpoint when the architecture predicts decay, consolidation, or delayed re-emergence of history effects.
8. Consequence-Block Test: Interpretation Without Experience-Relevance
A strong operational definition should distinguish interpretation from full experience-relevance. The consequence-block experiment permits a candidate state to be classified or locally interpreted but prevents it from altering memory, routing, correction, commitments, or Y. The immediate semantic representation may therefore exist while C_R and H_P are forced toward zero.
If the architecture's own experience-relevance measure still classifies the blocked state as fully experience-relevant, the measure is circular or insufficiently sensitive to consequence. If the classification correctly fails despite preserved local interpretation, the system demonstrates the intended distinction between 'represented now' and 'entered the receiver's consequential history.'
9. Minimal Heterogeneous Prototype for Confirmatory Testing
The first confirmatory platform should be large enough to instantiate heterogeneous roles but small enough to instrument completely. Six computational regions are sufficient as an engineering starting point, not as a theoretical requirement or privileged number.
Region | Primary role | Illustrative implementation |
|---|---|---|
B1 Interpretation | Construct candidate world-state interpretations from current input. | LLM or structured parser with uncertainty output. |
B2 Episodic memory | Retrieve and write consequential event records. | Vector/graph memory plus typed event ledger. |
B3 Simulation/planning | Generate counterfactual futures and candidate actions. | LLM planner, search process, or world-model rollout. |
B4 Salience/error | Estimate mismatch, contradiction, urgency, and resource cost. | Classifier, anomaly model, rule system, or small network. |
B5 Self/governance | Maintain receiver invariants, commitments, permissions, and local Y projections. | Deterministic state service plus policy rules; no language authority required. |
B6 Expression/action | Render or execute the currently committed outcome. | LLM renderer or task-specific actuator layer. |
The shared receiver state should be maintained in an instrumented state service rather than hidden only inside natural-language prompts. That service records Y, local projections, route weights, memory writes, unresolved alternatives, costs, conflicts, corrections, and fingerprints with timestamps. This creates an auditable causal trace for every confirmatory trial.
10. Comparator Architectures and Resource Matching
The full system cannot be judged against weak controls. At minimum, four architectures should be compared under matched base-model family, external information, context budget, and total compute where feasible:
Code | Comparator | Critical property absent |
|---|---|---|
C1 | Single persistent model/process | No heterogeneous multi-region receiver. |
C2 | Independent multi-agent ensemble with aggregation or voting | No recurrent shared receiver history; alternatives terminate at aggregation. |
C3 | Recurrent multi-agent system without persistent shared-Y governance | Recurrence exists, but no common history-bearing measuring architecture. |
C4 | Full Secretary Suite shared-Y receiver | Target architecture: shared governance, recurrence, unresolved field, consequential memory, projection coherence. |
Two fairness regimes are recommended. An equal-compute regime tests whether the proposed organization uses a fixed resource budget more effectively. An equal-capability regime allows C4 the recurrent compute required by its design and asks whether its claimed signatures are qualitatively different rather than merely stronger. Results should report both when practical.
11. Confirmatory Experimental Battery
Test | Primary endpoint | Predicted C4 signature | Decisive weakening result |
|---|---|---|---|
Semantic dissociation | S_sem | Meaning changes produce larger content-sensitive receiver divergence than matched surface-only changes. | No reliable difference after disruption and nuisance variables are matched. |
Projection fracture | Q⃗_Y / Q_Y* | Conflict becomes detectable and is repaired or produces systematic coherence loss. | Incompatible local Ys persist without correction or consequence. |
History swap | D_Y, H_P, D_hist | Matched present input is interpreted differently because prior consequence reconstructed Y/M. | History can be removed or swapped without changing later interpretation. |
Consequence block | C_R, H_P | Local interpretation survives while full experience-relevance classification fails. | Blocked states still meet the full criterion. |
Recurrence removal | C_R, Q⃗_Y, fingerprint | Global integration, repair, and receiver fingerprint weaken. | No meaningful change when recurrence is removed. |
Subconscious clamp | Reopening/metastability metrics | Forced convergence reduces delayed recovery, alternative preservation, and reopening. | Clamp produces no loss or improves all target signatures. |
Random-noise control | Promotion selectivity | Structured alternatives outperform equal-volume random variation in content-sensitive promotion. | Random variation reproduces the target signatures equally well. |
12. Preregistration Template
Every confirmatory experiment should be frozen in a preregistration record before the decisive data are generated. The record should be public or cryptographically time-stamped when publication strategy permits. At minimum it must specify:
The tested mechanism and directional hypothesis.
The exact architecture version, model versions, prompts or policies, memory schema, and Y schema.
Primary and secondary dependent variables, including the distance or similarity functions used.
The manipulated variable and the exact implementation of experimental, sham, and null conditions.
Primary time horizon k and any secondary horizons.
Seed or stochastic-replication policy and stopping rule.
Pilot/confirmatory split and a declaration that pilot trials will not be reclassified as confirmatory.
Thresholds or smallest effects of interest, including how they were derived.
Exclusion rules, failed-run handling, missing-data rules, and software/hardware faults that justify reruns.
Comparator resource budgets and whether the test is equal-compute, equal-latency, or equal-capability.
Primary statistical test, uncertainty interval, multiplicity rule, and decision criterion.
The negative result that will count against the mechanism rather than trigger a post-hoc redefinition.
13. Replication, Statistical Decision Logic, and Null Models
Because model behavior may be stochastic and non-Gaussian, paired experimental designs should be preferred wherever possible. The same receiver checkpoint and matched randomization schedule can be used across target and control conditions. A paired permutation test or bootstrap confidence interval is often appropriate when distributional assumptions are uncertain, provided the exact procedure is preregistered.
Sample size should not be chosen by convention alone. Pilot data should estimate baseline variability and the smallest theoretically meaningful effect. A simulation-based power analysis can then determine the number of independent receiver seeds, episodes, or perturbation pairs required for the desired detection probability. The chosen N and stopping rule must be fixed before confirmatory testing.
Null models should be mechanistically relevant. Recommended nulls include random-noise alternatives, shuffled semantic relations, sham Y perturbations, history records that are stored but denied causal access, and recurrent systems whose messages are exchanged without persistent shared governance. A model that beats only an obviously weaker baseline has not established the necessity of the proposed mechanism.
14. Primary Falsification Criteria
The full program should be considered weakened if any of the following survive replication under well-powered matched tests:
Semantic collapse: S_sem is not reliably greater for meaning changes than for meaning-preserving surface changes once disruption is matched.
Receiver federation: projection conflicts persist without detection, repair, uncertainty, or systematic receiver-level consequence.
History neutrality: consequential history can be removed, swapped, or neutralized without changing later interpretation in the predicted Y-dependent dimensions.
Consequence irrelevance: states remain classified as experience-relevant after their access to future routing, memory, correction, and Y update is blocked.
Recurrence dispensability: removing cross-region causal recurrence leaves the target integration, coherence, and fingerprint signatures unchanged.
Subconscious dispensability: structured unresolved alternatives can be eliminated without loss of reopening, counterfactual recovery, or metastable adaptation.
Comparator equivalence: matched simpler architectures reproduce the entire preregistered signature set with equal or greater simplicity.
No single successful test establishes phenomenal consciousness. Conversely, a failed operational test cannot be rescued by asserting that the system may still be conscious in an unmeasured sense. That would move the claim outside the scientific target defined by this program.
15. Implementation Sequence
Instrumentation build: implement typed Y, local projections, event ledger, route graph, memory writes, causal intervention hooks, and deterministic logging.
Reliability pilot: measure baseline variance and test-retest stability for D_Y, S_sem, C_R, H_P, and Q⃗_Y.
Semantic micro-world pilot: validate that meaning-changing and surface-only perturbations are matched for disruption and compute.
Coherence pilot: inject projection fractures and verify that the diagnostics can detect known conflicts.
Threshold freeze: define primary horizons, effect-size rules, null distributions, and confirmatory N from the pilot only.
Preregistered confirmatory battery: run C1–C4 plus planned ablations without changing metrics or thresholds.
Independent replication: rerun the frozen protocol with new seeds, new micro-world mappings, and ideally an independently implemented state service.
Report all outcomes: positive, negative, mixed, and null results, including tests that contradict the architecture.
16. What This Paper Adds — and What It Does Not
This paper adds measurement discipline to the Secretary Suite architecture. It operationalizes the four experience-relevance dimensions, supplies a grounded semantic-dissociation protocol, defines a receiver-coherence profile, formalizes projection-fracture and history-dependence tests, and specifies how confirmatory thresholds and null models should be frozen before decisive evaluation.
It does not claim that the proposed metrics are universal constants, that one composite score measures consciousness, that a particular number of regions is necessary, or that passing the battery proves phenomenal subjectivity. The metrics are implementation-specific instruments for testing a defined architectural claim. If a simpler system reproduces the same preregistered signatures, the additional Secretary Suite machinery loses explanatory necessity.
Conclusion
The computational-consciousness program has reached a point where additional conceptual expansion is less valuable than decisive measurement. The persistent receiver, structured unresolved field, dynamic present, shared Encoded Equilibrium, semantic consequence, and history-dependent reconstruction are now sufficiently specified to be subjected to controlled intervention.
The central scientific question is no longer whether the architecture sounds plausibly mind-like. It is whether a shared, recursively changing measuring architecture causes a distinctive pattern of semantic sensitivity, receiver coherence, causal consequence, and historical persistence that matched alternatives do not reproduce. The semantic-dissociation protocol tests meaning against surface statistics. The projection-fracture protocol tests whether heterogeneous regions genuinely form one receiver. The history-swap and consequence-block protocols test whether experience-relevant states enter and reconstruct the receiver's causal biography. The comparator battery tests whether those effects require the proposed architecture at all.
If the predicted signatures appear under preregistered conditions and degrade under targeted ablation, Secretary Suite will have moved from a conceptual architecture toward an experimentally supported computational research program. If the signatures fail, the theory has supplied a clear reason to revise or reject the relevant mechanism. Either outcome is scientifically useful.
Status of Claims
This paper is a Secretary Suite engineering and experimental-methodology extension built on the TSTOEAO-derived receiver architecture. Encoded Equilibrium, conditioned expression, gradients, boundaries, route structure, correction, cost, receiver conditions, and recursive state change are used as the theoretical measuring framework. The semantic-dissociation protocol, receiver-coherence profile, optional coherence composite, grounded semantic micro-world, projection-fracture protocol, normalized experience-relevance measurements, and preregistration procedure are proposed experimental constructions introduced for testing the architecture. None of these constructions establishes phenomenal consciousness, and no claim is made that a current commercial AI service already instantiates the proposed receiver.
References
Baars, Bernard J. 1988. A Cognitive Theory of Consciousness. Cambridge: Cambridge University Press.
Dehaene, Stanislas. 2014. Consciousness and the Brain: Deciphering How the Brain Codes Our Thoughts. New York: Viking.
Dennett, Daniel C. 1991. Consciousness Explained. Boston: Little, Brown and Company.
Nosek, Brian A., Charles R. Ebersole, Alexander C. DeHaven, and David T. Mellor. 2018. “The Preregistration Revolution.” Proceedings of the National Academy of Sciences 115 (11): 2600–2606.
Pearl, Judea. 2009. Causality: Models, Reasoning, and Inference. 2nd ed. Cambridge: Cambridge University Press.
Swygert, John. 2026a. “Computational Consciousness: A TSTOEAO- and EPH-Derived Architecture for Simulated Consciousness in Secretary Suite: From Structured Unresolved Possibility to Persistent Receiver Identity, Metastable Integration, and Recursive Becoming.” Secretary Suite / TSTOEAO Research Program.
Swygert, John. 2026b. “Computational Experience: Telemetry, Semantic Reality, and the Emergence of Consequential State in Simulated Consciousness.” Secretary Suite / TSTOEAO Research Program.