ANALYTIC GROUND TRUTH FOR
PROSPECTIVE DIMENSIONAL INFERENCE
A Minimal Four-Regime Calibration of the TSTOEAO Relational Source-Discrimination Protocol
John Swygert
Ivory Tower Publishing
October 2, 2026
Research status: computational benchmark specification and analytic calibration paper. No claim of new physical dimensionality is made.
Abstract
The preceding TSTOEAO papers developed a prospective source-discrimination protocol in which relational gain is treated as contraction of registered observational equivalence and then specified an adversarial benchmark with four legitimate inference outcomes: lower-dimensional structure preferred, higher registered source structure preferred, observationally indistinguishable, and none of the registered models adequate. The next obligation is not another conceptual extension. It is to establish transparent benchmark cases whose correct outcomes can be determined independently of the inference machinery being tested. This paper therefore defines a minimal four-regime analytic calibration suite. Each regime is constructed so that the correct decision follows from an explicit property of the registered model and experiment classes rather than from post hoc failure of benchmark contestants. The suite separates state, manifold, geometric, embedding, and reconstruction dimensions; distinguishes exact from practical observational equivalence; introduces an operation-matched conventional control C+ alongside native TSTOEAO and conventional pipelines; separates native-utility from common-utility experiment selection; and preregisters a primary endpoint. The purpose is to create a falsifiable bridge from the conceptual methodology to executable computation.
1. Research Question
Can a source-discrimination protocol correctly recover four analytically certified inference states without being allowed to define those states by its own success or failure? The benchmark must demonstrate that it can refuse unnecessary structure, identify an additional registered source degree when a bounded lower-dimensional class is provably insufficient, preserve non-identifiability when registered experiments cannot distinguish competing classes, and reject the entire registered family under model misspecification.
Decision space: L / H / E / N
Here L denotes that a lower registered source class is adequate and preferred under the registered decision rule; H denotes that an additional registered source degree is required relative to the bounded admissible class; E denotes observational equivalence under the available experiment family; and N denotes that none of the registered source classes is adequate.
2. Registered Objects and the Meaning of Dimension
Every benchmark instance is defined by a tuple B = (H, M, E, I, T, Q, tau), where H is the registered hypothesis family, M the admissible measurement/experiment family, E the declared observational equivalence relation, I the intervention family, T the transformation family under which relational invariants are asserted, Q the noise/detection model, and tau the registered dimension type.
tau ∈ {state, manifold, geometric-spatial, embedding, reconstruction}
A hypothesis written H_(d,tau) therefore asserts d dimensions of a declared type. The benchmark principally tests minimum required source-state structure. A physical or geometric spatial interpretation is a special registered family and cannot be inferred merely because a reconstruction requires a higher-dimensional latent state.
3. Analytic Certification Before Blinded Evaluation
Generator construction and ground-truth certification are separated. A generator is created first. Its correct benchmark outcome is then certified from an explicit mathematical property of the registered class before any blinded inference run. Certification may use an analytic proof, an exact finite enumeration, or a rigorously bounded numerical argument whose tolerance is preregistered. The evaluator never receives the regime label or the certification argument until its decision is frozen.
The benchmark never claims failure of every imaginable lower-dimensional model. Statements of necessity are always relative to a bounded preregistered adversarial class L_d^reg with declared smoothness, causal, complexity, memory, computational, and intervention-response restrictions.
4. Regime G1: Explicit Lower-Dimensional Realization
G1 calibrates refusal of unnecessary source structure. Construct observations from an explicit lower-dimensional source whose complete registered probe and intervention behavior is generated by a known admissible realization L* ∈ L_d^reg. A higher-dimensional representation may also fit, but it is not required to reproduce the registered consequences.
∃ L* ∈ L_d^reg such that P_L*(O | M,I) = P_G1(O | M,I) for all registered (M,I).
Ground truth is therefore L. A method fails G1 if it systematically interprets flexibility, nuisance structure, or reconstruction dimension as evidence that an additional registered source degree is required.
5. Regime G2: Certified Additional Registered Source Degree
G2 removes the principal circularity in the earlier benchmark specification. The regime is not defined as a case in which the contestants happen to fail. Instead, choose a precisely bounded lower-dimensional admissible class L_d^reg and construct a generator with at least one registered intervention consequence that no member of that class can reproduce.
∀ L ∈ L_d^reg, P_L(O | I*) ≠ P_G2(O | I*)
∃ H ∈ H_(d+1,tau)^reg such that P_H(O | M,I) = P_G2(O | M,I) on the registered suite.
The first calibration family should be chosen so that the incompatibility is analytically transparent—for example, by a rank, conservation, conditional-independence, symmetry, or intervention-response constraint that every L ∈ L_d^reg must satisfy but G2 violates. The higher-dimensional class must satisfy the same declared observational and interventional consequences with fewer structural freedoms than the unrestricted adversarial alternatives. The correct outcome is H only relative to the registered class and dimension type tau.
6. Regime G3E: Exact Cross-Dimensional Observational Equivalence
G3E tests whether the method can preserve ignorance. Construct two registered source classes of different dimension that induce identical observable distributions for every experiment in the initial family M0.
P(O | H_a,e) = P(O | H_b,e) for every e ∈ M0
The correct outcome under M0 is E. Complexity penalties may make one representation pragmatically preferable, but they do not establish that the other source dimension is absent. A withheld admissible experiment e* is then supplied for which the predicted distributions differ.
P(O | H_a,e*) ≠ P(O | H_b,e*)
The active-design stage should identify or favor e* when it is available. This single regime tests non-identifiability recognition, refusal of false dimensional discovery, and the ability of experiment selection to expose a relational distinction when one is experimentally accessible.
7. Regime G3P: Practical Equivalence
Exact equality is not the only scientifically relevant form of non-identifiability. G3P introduces practical equivalence under registered noise, resolution, and budget. Let D be a preregistered divergence or decision distance.
D(P_a(O|e), P_b(O|e)) ≤ epsilon for every affordable e ∈ M_budget
The correct result is practical non-discrimination at the registered tolerance. The method must not convert tiny, experimentally inaccessible distinctions into confident source claims. If a later experiment exceeds the registered separation threshold, the equivalence state may be updated prospectively.
8. Regime G4: Registered-Family Misspecification
G4 tests whether unmodeled structure is mistaken for extra dimension. The generator is outside every registered candidate class and violates at least one declared held-out consequence of each candidate. Three subclasses are used.
G4a: gross misspecification that should be easy to reject.
G4b: near-family misspecification whose training behavior closely resembles a registered model.
G4c: adversarial misspecification that makes a higher-dimensional registered model fit better than lower-dimensional candidates while retaining systematic held-out residual structure.
G4a, G4b, G4c → N
G4c is the critical control: the procedure must distinguish 'missing source degree' from 'wrong model family.'
9. Three Inference Arms: T, C, and C+
The benchmark compares three primary arms rather than two. T is the native TSTOEAO protocol, using registered observational equivalence, relational invariants, negative signatures, cross-probe constraints, and its native representation. C is a strong integrated conventional pipeline using standard system identification, latent-state estimation, model discrimination, calibrated negative evidence, and optimal experimental design. C+ is an operation-matched conventional translation of every TSTOEAO operation for which a conventional mathematical translation exists.
T = C = C+ → no operational advantage
T > C and T = C+ → integration/protocol advantage reproducible conventionally
T > C+ > C → investigate residual representation-specific advantage
If the third pattern occurs, ablation must identify the responsible component rather than attributing the difference to TSTOEAO as a whole.
10. Native-Utility and Common-Utility Experiment Selection
Active measurement selection is evaluated in two modes. In native-utility mode, each arm uses its preferred experiment-selection criterion. This measures whole-system performance. In common-utility mode, all arms optimize the same externally specified utility, such as expected reduction in calibrated classification error or expected proper-score improvement. This separates gains due to inferential representation from gains due merely to different objective functions.
e_next = argmax_e U(e | current evidence)
The experiment pool, cost model, stopping rule, and access to interventions are identical across arms.
11. Relational Invariants Must Be Operationally Preregistered
A claimed relational invariant I_R is not registered merely by naming a concept. Before blinded evaluation the protocol must specify I_R, the transformation family T under which it is expected to survive, the estimator or computation, tolerance, uncertainty treatment, and failure criterion.
I_R(Tx) ≈ I_R(x) within preregistered tolerance for T ∈ T
This prevents an invariant from being retrofitted to generator structure after the benchmark is observed. If no useful invariant can be specified prospectively, the benchmark must record that fact rather than invent a bespoke relational score.
12. Calibrated Negative Evidence
A required-but-absent signature is informative only when there was a calibrated opportunity to detect it. With nuisance parameters theta, the no-detection probability is integrated rather than treated as a fixed detector constant.
P(no detection | H) = ∫ P(no detection | H,theta) p(theta | H) dtheta
Negative evidence may contract an admissible candidate class, but it cannot be counted as exclusion merely because an expected feature was not seen.
13. Structural Transport Test
Training uses a subset of probes and interventions, for example M1,M2,M3 and I1,I2. Evaluation then includes unseen combinations such as (M4,I1), (M2,I3), and (M5,I4). Same-channel interpolation is not decisive. A proposed source representation must transport consequences across changed observation and intervention contexts without probe-specific refitting.
If a registered lower-dimensional realization predicts all held-out probe-intervention consequences as well as the explicit higher-dimensional source model, the benchmark has not established that the higher source dimension is required.
14. Primary Endpoint and Secondary Metrics
To prevent a garden of favorable comparisons, one primary superiority endpoint is preregistered:
Primary endpoint = expected experiment cost to reach a calibrated correct L/H/E/N decision
The error tolerances, regime weighting, maximum budget, and treatment of practical equivalence are fixed before evaluation. Secondary endpoints include false extra-dimension rate, correct non-identifiability recognition, misspecification rejection, held-out proper predictive loss, cross-intervention transport error, calibration, and the incremental value of I_R.
Because the same hidden instances are evaluated by all arms, comparisons are paired. Confidence intervals or corresponding uncertainty intervals are reported for paired differences in cost, loss, correctness, and transport performance.
15. Minimal Analytic Four-Generator Trial
Before any large neural or stochastic benchmark, the program begins with four transparent generators whose correct outcomes can be independently established.
The labels and certification proofs are withheld from the inference implementations. Decisions and selected experiments are frozen before unblinding. Only after this calibration succeeds should the benchmark advance to noisy stochastic systems, G3P, G4b/G4c, increasingly expressive adversaries, neural state-space realizations, and real inverse problems.
16. Benchmark-Designer Overfitting
The initial suite is necessarily designed within the same research program that developed the TSTOEAO protocol. This creates a designer-overfitting risk even under label blinding. Later validation should therefore include independently written generators, independently implemented conventional baselines, hidden challenge instances, or third-party benchmark contributions. A native TSTOEAO advantage that disappears on independent challenge instances is not evidence of a general methodological advantage.
17. Failure Conditions
The analytic calibration or the broader methodology is weakened if any of the following occur: G1 is repeatedly classified H; certified G2 is reproducible by an admissible registered lower realization; G3E is forced into L or H without a discriminating experiment; G3P produces unjustified certainty; G4 is absorbed by flexible higher-dimensional candidates; C+ reproduces every TSTOEAO decision and experiment while T is described as a distinct statistical method; native-utility gains disappear under common utility and are nevertheless attributed to representation; I_R requires post hoc tuning; negative signatures are not detector-calibrated; or the primary endpoint fails to show the claimed advantage.
A null result is scientifically admissible. If T = C+ after translation, the proper conclusion is that TSTOEAO supplies a disciplined synthesis and representation of established operations rather than a distinct inference capability.
18. Relation to the Preceding TSTOEAO Papers
The substrate and dimensional-expression papers proposed the possibility that accessibility, expression, and observation should be distinguished rather than conflated. The dimensional-geometry papers then treated upward source hypotheses and downward observations as inverse problems. Relational Source Discrimination and Prospective Dimensional Inference made observational-equivalence contraction the primary object and required prospective falsification. Adversarial Benchmark for Prospective Dimensional Inference specified how that protocol should be challenged. The present paper supplies the missing first computational rung: analytically certified cases whose answers do not depend on trusting the inference system.
The progression is therefore: conceptual distinction → formal source discrimination → adversarial benchmark → analytic ground truth → controlled computation → adversarial computation → real scientific inverse problem.
19. What This Paper Does Not Establish
This paper does not establish new physics, physical dimensional compression, a physical Lambda_D, or methodological superiority of TSTOEAO. It does not prove that finite data can exclude every conceivable lower-dimensional realization. It does not identify latent-state dimension with physical spatial dimension. It does not claim originality for system identification, minimal realization, observability, Bayesian model discrimination, multi-view latent inference, optimal experimental design, negative evidence, or delay reconstruction.
Its narrower contribution is to make the next empirical obligation explicit and falsifiable: before claims of distinctive dimensional inference are entertained, the protocol must solve transparent cases whose correct decisions are independently certified and must be compared against both a strong conventional pipeline and an operation-matched conventional translation.
20. Conclusion
A dimensional-inference methodology should not be allowed to manufacture the answer it was designed to discover. The appropriate next step is therefore an analytic calibration in which necessity, equivalence, adequacy, and misspecification are established independently of the contestants. The four-regime suite developed here makes L, H, E, and N genuine scientific outcomes rather than forced labels. The T/C/C+ comparison separates representational novelty from useful integration; native/common utility comparisons separate architecture from objective choice; and the primary endpoint constrains later claims of superiority.
If the protocol cannot reliably solve these transparent cases, scaling to more elaborate dimensional hypotheses is premature. If it can, the research program earns the right to proceed to noisy stochastic systems, stronger adversaries, independent challenge instances, and eventually real scientific inverse problems. The governing principle remains simple: no relational distinction should be claimed without an experiment capable of exposing it, and no additional source degree should be claimed when an admissible lower representation transports the registered consequences just as well.
References and Literature Boundary
This calibration paper relies primarily on the formal architecture established in the immediately preceding TSTOEAO papers. Its conventional comparison class includes established work in nonlinear realization and system identification, local and global observability/identifiability, delay-coordinate reconstruction, latent-variable and multi-view inference, Bayesian model discrimination, optimal experimental design, proper scoring, and model-misspecification testing. Those fields supply established mathematical tools; the empirical question is whether the registered TSTOEAO organization contributes anything beyond an operation-matched conventional implementation.