Tuesday, August 25, 2026

Experiments A, B, and C: Experiment A Relational Codec vs. Brotli, Zstandard, Deduplication, and Delta Encoding: A Secretary Suite Project

Experiments A, B, and C:

Experiment A Relational Codec vs. Brotli, Zstandard, Deduplication, and Delta Encoding: A Secretary Suite Project 

DOI: [To be assigned]

John Swygert

August 25, 2026

Abstract

This experimental protocol directly tests the Relational Codec against established compression and synchronization baselines. It is designed to prevent an easy but meaningless victory over uncompressed data. The experiment uses identical corpora, exact reconstruction, frozen preprocessing, and a range of shared-state conditions. It measures physical bytes, metadata overhead, encoding and decoding time, memory, failure recovery, and total cost. The key hypothesis is conditional rather than universal: relational reference and delta methods should outperform conventional full-object compression primarily when sender and receiver share substantial reusable state, while Brotli or Zstandard may remain superior for novel independent objects.

Governing principle: minimize total storage or transmission cost subject to explicit reconstruction, integrity, provenance, and reliability requirements.

1. Hypotheses

H1: At low shared-state overlap, Brotli or Zstandard will match or outperform relational encoding after metadata overhead is included.

H2: At high overlap, deduplication and delta methods will reduce transfer volume substantially.

H3: A hybrid relational router will approach the best available method across both regimes without excessive decision overhead.

2. Corpus

Use at least six workload families: repeated prose, evolving prose versions, code repository snapshots, templated structured reports, mixed knowledge shards, and random/high-entropy controls.

Include AInunnaki only as one prose workload; do not tune parameters exclusively to it.

3. Systems Compared

RAW, Brotli, Zstandard, content-defined chunk deduplication, generic delta encoding, Relational Codec shard-reference mode, shard+delta mode, and adaptive hybrid mode.

4. Shared-State Conditions

Test receiver overlap at 0%, 10%, 25%, 50%, 75%, 90%, 99%, and 100% of reusable base content. Verify shared-state manifests before each run.

5. Exactness

All primary comparisons are lossless. Reconstructed artifacts must match cryptographic hashes of originals.

6. Measurements

Original bytes, transmitted bytes, stored unique bytes, manifest bytes, encoder time, decoder time, peak memory, CPU time, cache lookup time, dependency depth, and fallback bytes.

7. Derived Metrics

Transmission Ratio=TxBytes/OriginalBytes. Storage Ratio=UniqueStored/TotalLogicalBytes. Savings=(Baseline-Method)/Baseline. EndToEndCost uses disclosed application weights.

8. Procedure

Freeze preprocessing; build sender corpus; create receiver state according to overlap condition; encode; transmit representation; reconstruct; verify hash; repeat; then rotate workload and seed.

9. Failure Injection

Remove 1%, 5%, and 10% of referenced cached shards without telling the encoder. The decoder must detect missing state and recover through targeted fallback rather than silently producing incorrect output.

10. Analysis

Compare paired results within each workload and overlap level. Identify crossover points where relational methods become cheaper than conventional compression.

11. Falsification

If the relational system does not beat or closely track the best baseline in the high-overlap regimes for which it is designed, revise or reject the mechanism.

12. Deliverables

Publish raw benchmark inputs, frozen configuration, manifests, exact software versions, per-run metrics, and negative results.

13. Conclusion

This experiment answers the central practical question directly: under what shared-state conditions, if any, does a relational codec reduce real storage or network cost beyond established methods?

Methodological Guardrails

  • Compare against strong existing baselines; never claim gains relative only to raw/uncompressed data.

  • Count manifests, hashes, provenance, indices, repair traffic, and routing overhead as real cost.

  • Keep logical shard boundaries separate from physical storage/chunk boundaries.

  • Distinguish exact byte reconstruction from functional or semantic reconstruction.

  • Treat similarity as evidence of resemblance, not proof of derivation or shared provenance.

  • Publish crossover points and negative results where conventional methods win.

  • Use cryptographic integrity checks for exact reconstruction experiments.

  • Treat security, privacy, and access boundaries as constraints, not optional afterthoughts.

  • Mappability to TSTOEAO is not validation; empirical advantage must be demonstrated.

References

Shannon, C. E. (1948). A Mathematical Theory of Communication. Bell System Technical Journal, 27, 379-423, 623-656.

Deutsch, P. (1996). DEFLATE Compressed Data Format Specification version 1.3. RFC 1951.

Korn, D. G., MacDonald, J. P., Mogul, J. C., & Vo, K.-P. (2002). The VCDIFF Generic Differencing and Compression Data Format. RFC 3284.

Tridgetgell, A., & Mackerras, P. (1996). The rsync algorithm. Australian National University Technical Report TR-CS-96-05.

Alakuijala, J., & Szabadka, Z. (2016). Brotli Compressed Data Format. RFC 7932.

Collet, Y., & Kucherawy, M. (2018). Zstandard Compression and the application/zstd media type. RFC 8878.

Xia, W., Jiang, H., Feng, D., Douglis, F., Shilane, P., Hua, Y., Fu, M., Zhang, Y., & Zhou, Y. (2016). FastCDC: a Fast and Efficient Content-Defined Chunking Approach for Data Deduplication. USENIX Annual Technical Conference.

Swygert, J. (2026). 200 From Language to Computation: Linguistics, Punctuation, Mathematics, and Programming as a Unified Relational Architecture in Large Language Models. Ivory Tower Publishing.

Swygert, J. (2026). 203 The Shard as a Relational Unit: Deconstruction, Reconstruction, and Knowledge Architecture. Ivory Tower Publishing.

Swygert, J. (2026). 204 The Statistical Shard Library: Measuring Retrieval, Reuse, Frequency, and Relational Importance. Ivory Tower Publishing.

Swygert, J. (2026). 205 Symbolic Fingerprints: Provenance, Reuse Detection, and Plagiarism Prevention through Relational Analysis. Ivory Tower Publishing.

Swygert, J. (2026). 206 Relational Reconstruction: Building Novel Outputs from Provenance-Aware Knowledge Shards. Ivory Tower Publishing.

Swygert, J. (2026). 207 The Relational Intelligence Engine: From Language Structure to Adaptive Machine Reasoning. Ivory Tower Publishing.

Swygert, J. (2026). 208 The Empirical Relational Benchmark: A Hybrid Character, Word, Shard, and Provenance Test Architecture for Language Models and Knowledge Systems. Ivory Tower Publishing.


Experiment B Shared Shard Library Transmission Test: Minimum Data Required for Exact Reconstruction: A Secretary Suite Project

DOI: [To be assigned]

John Swygert

August 25, 2026

Abstract

This protocol measures the minimum amount of new data that must cross a network boundary when two endpoints share a verified Shard Library. Rather than comparing only compressors, it decomposes each target into receiver-known shards, changed shards, novel shards, relational changes, and metadata changes. The experiment progressively increases shared knowledge and records the payload needed for exact reconstruction. It thereby tests the project's original transmission objective in its most direct form: transmit less data from server to receiver by exploiting shared verified state.

Governing principle: minimize total storage or transmission cost subject to explicit reconstruction, integrity, provenance, and reliability requirements.

1. Research Question

Given target artifact D and receiver state K, what is the smallest practical payload M such that Decode(M,K)=D exactly?

2. Experimental Model

Sender and receiver begin from a common library version, then diverge. The sender creates or receives a target artifact containing mixtures of old, modified, reordered, and novel material.

3. Payload Classes

Known-shard references, changed-shard deltas, novel payload, relation edits, ordering instructions, metadata/provenance updates, integrity hashes, and fallback payload.

4. Shared-State Sweep

Construct receiver libraries at controlled overlap levels. Overlap must be measured both in bytes and in logical shards because the two can diverge.

5. Reconstruction Manifest

The sender generates a manifest identifying base state, required shards, relations, deltas, novel data, expected artifact hash, and optional provenance graph hash.

6. Exactness Gate

A run succeeds only when the receiver reconstructs the exact target and matches the artifact hash. Missing shared state must trigger a repair request.

7. Minimum-Payload Search

For each target, compare candidate representations and choose the smallest valid payload after including all required metadata. Do not exclude manifest cost.

8. Order-Only Cases

Create targets whose content is unchanged but whose shard order differs. Measure whether permutation instructions beat retransmission.

9. Relation-Only Cases

Create knowledge targets whose content is unchanged but provenance or dependency edges change. Measure payload for graph-only reconstruction.

10. Local Modification Cases

Modify a small region within a large shard and compare whole-shard replacement, byte delta, token delta, and sub-sharding.

11. Novelty Cases

Increase genuinely new material until relational referencing stops helping. The point at which conventional compression wins is a key empirical boundary.

12. Results Table

For every run report target bytes, receiver-known bytes, known shard fraction, manifest bytes, new payload bytes, repair bytes, total transmitted bytes, decode time, and exactness result.

13. Primary Outcome

Define Minimum Transmission Fraction MTF = TotalTransmittedBytes / TargetBytes. Lower is better, but only for exact successful reconstruction.

14. Secondary Outcome

Define Shared-State Leverage L = (BaselineTx - HybridTx) / VerifiedSharedBytes to estimate how effectively prior shared state is converted into network savings.

15. Falsification

If manifest and lookup overhead eliminate gains except in trivial duplicate cases, the broader relational transmission claim must be narrowed.

16. Conclusion

This experiment operationalizes the question 'How little must cross the boundary?' and provides a direct measure of whether the Shard Library is moving toward reduced network volume.

Methodological Guardrails

  • Compare against strong existing baselines; never claim gains relative only to raw/uncompressed data.

  • Count manifests, hashes, provenance, indices, repair traffic, and routing overhead as real cost.

  • Keep logical shard boundaries separate from physical storage/chunk boundaries.

  • Distinguish exact byte reconstruction from functional or semantic reconstruction.

  • Treat similarity as evidence of resemblance, not proof of derivation or shared provenance.

  • Publish crossover points and negative results where conventional methods win.

  • Use cryptographic integrity checks for exact reconstruction experiments.

  • Treat security, privacy, and access boundaries as constraints, not optional afterthoughts.

  • Mappability to TSTOEAO is not validation; empirical advantage must be demonstrated.

References

Shannon, C. E. (1948). A Mathematical Theory of Communication. Bell System Technical Journal, 27, 379-423, 623-656.

Deutsch, P. (1996). DEFLATE Compressed Data Format Specification version 1.3. RFC 1951.

Korn, D. G., MacDonald, J. P., Mogul, J. C., & Vo, K.-P. (2002). The VCDIFF Generic Differencing and Compression Data Format. RFC 3284.

Tridgetgell, A., & Mackerras, P. (1996). The rsync algorithm. Australian National University Technical Report TR-CS-96-05.

Alakuijala, J., & Szabadka, Z. (2016). Brotli Compressed Data Format. RFC 7932.

Collet, Y., & Kucherawy, M. (2018). Zstandard Compression and the application/zstd media type. RFC 8878.

Xia, W., Jiang, H., Feng, D., Douglis, F., Shilane, P., Hua, Y., Fu, M., Zhang, Y., & Zhou, Y. (2016). FastCDC: a Fast and Efficient Content-Defined Chunking Approach for Data Deduplication. USENIX Annual Technical Conference.

Swygert, J. (2026). 200 From Language to Computation: Linguistics, Punctuation, Mathematics, and Programming as a Unified Relational Architecture in Large Language Models. Ivory Tower Publishing.

Swygert, J. (2026). 203 The Shard as a Relational Unit: Deconstruction, Reconstruction, and Knowledge Architecture. Ivory Tower Publishing.

Swygert, J. (2026). 204 The Statistical Shard Library: Measuring Retrieval, Reuse, Frequency, and Relational Importance. Ivory Tower Publishing.

Swygert, J. (2026). 205 Symbolic Fingerprints: Provenance, Reuse Detection, and Plagiarism Prevention through Relational Analysis. Ivory Tower Publishing.

Swygert, J. (2026). 206 Relational Reconstruction: Building Novel Outputs from Provenance-Aware Knowledge Shards. Ivory Tower Publishing.

Swygert, J. (2026). 207 The Relational Intelligence Engine: From Language Structure to Adaptive Machine Reasoning. Ivory Tower Publishing.

Swygert, J. (2026). 208 The Empirical Relational Benchmark: A Hybrid Character, Word, Shard, and Provenance Test Architecture for Language Models and Knowledge Systems. Ivory Tower Publishing.


Experiment C Adaptive Hybrid Routing: Selecting the Lowest-Cost Representation for Each Data Class: A Secretary Suite Project

DOI: [To be assigned]

John Swygert

August 25, 2026

Abstract

This protocol tests whether an adaptive router can choose among raw transfer, conventional compression, deduplication, delta encoding, shard referencing, and relational-template reconstruction more efficiently than any one fixed method across heterogeneous workloads. The experiment separates oracle performance from learnable routing performance. First, every candidate method is run so the true best choice is known. A router then attempts to predict that choice from inexpensive features such as file size, entropy estimate, shared-state overlap, chunk reuse, edit distance, shard reuse, and template availability. The core metric is regret: the additional total cost incurred relative to the oracle.

Governing principle: minimize total storage or transmission cost subject to explicit reconstruction, integrity, provenance, and reliability requirements.

1. Objective

The hybrid system should not maximize sophistication. It should minimize cost by choosing the simplest method adequate for each object or region.

2. Candidate Methods

RAW, Brotli, Zstandard, deduplication, delta, shard-reference, shard+delta, relational-template, and selected combinations.

3. Oracle

Run all methods on each sample. The method with lowest declared total cost becomes the oracle label for that sample.

4. Routing Features

Original size, estimated entropy, MIME/type class, similarity to prior object, receiver overlap, chunk hit rate, shard hit rate, edit distance estimate, expected manifest size, and template match score.

5. Router Models

Begin with transparent rules and logistic/tree models before deep models. Complexity must earn its cost.

6. Budget Constraint

Selection can be expressed as argmax over expected information savings subject to latency or compute budget. The router may stop after a cheap stage if confidence is high.

7. Hierarchical Routing

Stage 1 decides whether shared-state methods are plausible. Stage 2 chooses byte/chunk/shard scale. Stage 3 selects compression and delta parameters. Stage 4 validates expected savings before commitment.

8. Metrics

Top-1 method-selection accuracy, regret versus oracle, total bytes, total CPU, latency, memory, fallback rate, and routing overhead.

9. Region-Level Routing

A single file may contain compressible and incompressible regions. An advanced condition allows the router to partition an object and use different representations per region.

10. Robustness

Test distribution shift: train routing rules on prose and code, then introduce images or binary assets. The router should fall back safely rather than hallucinate relational savings.

11. Falsification

If routing overhead plus mistakes erase the gains from specialization, fixed methods are preferable. Report that outcome.

12. Deployment Criterion

Deploy adaptive routing only when average regret is low and its savings exceed profiling and model overhead across the intended workload.

13. Conclusion

The adaptive hybrid experiment tests the project's plural principle directly: no representation is privileged; the system should use whichever valid method reconstructs the target at the lowest verified cost.

Methodological Guardrails

  • Compare against strong existing baselines; never claim gains relative only to raw/uncompressed data.

  • Count manifests, hashes, provenance, indices, repair traffic, and routing overhead as real cost.

  • Keep logical shard boundaries separate from physical storage/chunk boundaries.

  • Distinguish exact byte reconstruction from functional or semantic reconstruction.

  • Treat similarity as evidence of resemblance, not proof of derivation or shared provenance.

  • Publish crossover points and negative results where conventional methods win.

  • Use cryptographic integrity checks for exact reconstruction experiments.

  • Treat security, privacy, and access boundaries as constraints, not optional afterthoughts.

  • Mappability to TSTOEAO is not validation; empirical advantage must be demonstrated.

References

Shannon, C. E. (1948). A Mathematical Theory of Communication. Bell System Technical Journal, 27, 379-423, 623-656.

Deutsch, P. (1996). DEFLATE Compressed Data Format Specification version 1.3. RFC 1951.

Korn, D. G., MacDonald, J. P., Mogul, J. C., & Vo, K.-P. (2002). The VCDIFF Generic Differencing and Compression Data Format. RFC 3284.

Tridgetgell, A., & Mackerras, P. (1996). The rsync algorithm. Australian National University Technical Report TR-CS-96-05.

Alakuijala, J., & Szabadka, Z. (2016). Brotli Compressed Data Format. RFC 7932.

Collet, Y., & Kucherawy, M. (2018). Zstandard Compression and the application/zstd media type. RFC 8878.

Xia, W., Jiang, H., Feng, D., Douglis, F., Shilane, P., Hua, Y., Fu, M., Zhang, Y., & Zhou, Y. (2016). FastCDC: a Fast and Efficient Content-Defined Chunking Approach for Data Deduplication. USENIX Annual Technical Conference.

Swygert, J. (2026). 200 From Language to Computation: Linguistics, Punctuation, Mathematics, and Programming as a Unified Relational Architecture in Large Language Models. Ivory Tower Publishing.

Swygert, J. (2026). 203 The Shard as a Relational Unit: Deconstruction, Reconstruction, and Knowledge Architecture. Ivory Tower Publishing.

Swygert, J. (2026). 204 The Statistical Shard Library: Measuring Retrieval, Reuse, Frequency, and Relational Importance. Ivory Tower Publishing.

Swygert, J. (2026). 205 Symbolic Fingerprints: Provenance, Reuse Detection, and Plagiarism Prevention through Relational Analysis. Ivory Tower Publishing.

Swygert, J. (2026). 206 Relational Reconstruction: Building Novel Outputs from Provenance-Aware Knowledge Shards. Ivory Tower Publishing.

Swygert, J. (2026). 207 The Relational Intelligence Engine: From Language Structure to Adaptive Machine Reasoning. Ivory Tower Publishing.

Swygert, J. (2026). 208 The Empirical Relational Benchmark: A Hybrid Character, Word, Shard, and Provenance Test Architecture for Language Models and Knowledge Systems. Ivory Tower Publishing.

No comments:

Post a Comment