Thursday, September 10, 2026

The Provenance of Discovery: Intellectual Chain of Custody for AI-Assisted Mathematics: A Secretary Suite Project

The Provenance of Discovery: Intellectual Chain of Custody for AI-Assisted Mathematics

A Secretary Suite Project

John Swygert
Ivory Tower Publishing
September 10, 2026


Abstract

Artificial intelligence is beginning to participate directly in mathematical discovery rather than functioning merely as a calculator, search interface, or writing assistant. Systems can now explore conjectures, construct arguments, coordinate large populations of reasoning agents, formalize proofs, search existing literature, evaluate intermediate hypotheses, and in some cases produce mathematical results whose immediate intellectual origin is principally computational. OpenAI's 2026 announcements of substantial progress on multiple longstanding mathematical problems, followed by its announced solution to the Navier–Stokes Millennium Prize problem, demonstrate how quickly this transition is occurring. OpenAI reported that an internal version of its Astra model generated ten significant mathematical and theoretical-computer-science results and subsequently formalized the arguments in Lean. In September 2026, the company reported a much larger multi-agent effort using approximately 10,000 concurrent agents in the group that produced its Navier–Stokes result.

These developments have simultaneously produced disputes over intellectual priority, training data, user interactions, unpublished research, attribution, and the possibility that mathematical conversations conducted through AI systems might later contribute to improved models. Such concerns are understandable, but the present debate frequently collapses several different questions into one: whether information was legally or contractually permitted to be used, whether it actually entered a training or retrieval pipeline, whether a later system was technically exposed to it, whether it causally influenced a particular discovery, and whether scholarly attribution is therefore warranted.

These questions cannot be resolved by suspicion, similarity, institutional assertion, or retrospective memory. They require provenance.

This paper proposes an intellectual chain-of-custody architecture for AI-assisted mathematics. It extends the Personal Provenance Ledger framework developed through Secretary Suite into mathematical discovery and model development, establishing a persistent record connecting human contributors, AI interactions, source materials, permissions, training eligibility, datasets, model checkpoints, computational activities, intermediate discoveries, formal verification, and publication.

The central principle is simple:

Attribution should follow demonstrable intellectual lineage.

Neither an AI laboratory nor an individual researcher should be required to prove or defend intellectual descent through speculation when technical infrastructure can instead preserve evidence. The objective is not to restrict mathematical exchange. It is to make increasingly collaborative human–machine discovery more open, more trustworthy, more reproducible, and ultimately more productive.

Keywords: artificial intelligence, mathematics, provenance, intellectual provenance, chain of custody, AI-assisted discovery, training data, attribution, scholarly credit, model provenance, research integrity, formal verification, Secretary Suite


Introduction

Mathematics has always been cumulative.

Every theorem exists within an intellectual landscape shaped by earlier definitions, techniques, failed approaches, conjectures, proofs, conversations, correspondence, lectures, manuscripts, and communities of researchers. Even a radically original proof rarely emerges without inheritance from a larger mathematical tradition.

Artificial intelligence changes neither this fact nor the fundamental nature of mathematics.

What AI changes is the scale, velocity, opacity, and complexity through which mathematical inheritance may occur.

A human mathematician may remember reading a paper twenty years earlier. Another may derive an approach after hearing a seminar. A third may discover something independently. Traditional scholarly practice handles these ambiguities imperfectly through citations, publication dates, correspondence, notebooks, preprints, and professional judgment.

AI introduces vastly more complicated pathways.

A mathematical concept might appear in a published article, an arXiv manuscript, an online discussion, a private conversation with an AI assistant, an evaluation dataset, a retrieved document, a fine-tuning corpus, an intermediate synthetic dataset, an agent-generated result, or a later model checkpoint. A model may reproduce an idea because it encountered the original. It may independently rediscover the same structure. It may combine several unrelated sources into something new. It may be influenced indirectly through model improvement without anyone involved in a later inference ever seeing the original interaction.

Without provenance, these possibilities can become nearly impossible to distinguish.

The problem is already visible.

In September 2026, mathematician Andreas Thom publicly questioned whether conversations he or colleagues had with ChatGPT could have contributed to later OpenAI mathematical work involving non-sofic groups. Reporting by The Verge described Thom as seeking evidence regarding whether his interactions entered training data or otherwise influenced the resulting system. OpenAI had previously stated regarding its Navier–Stokes work that its researchers and agents did not access specific unpublished user data, while adding that it could not completely rule out the possibility that de-identified data derived from product usage had indirectly contributed to improving its models.

That distinction is technically important.

It is also precisely the kind of distinction that ordinary scholarly infrastructure was never designed to resolve.

A new infrastructure is therefore required.


Mathematics Should Be the Primary Concern

The first principle of this discussion should be easy to state:

The advancement of mathematics matters more than institutional ego.

If an artificial intelligence system discovers a valid proof to a longstanding mathematical problem, the appropriate first response should be intellectual curiosity.

Is the proof correct?

What new ideas does it contain?

Can those ideas generalize?

Do they expose structures human mathematicians had overlooked?

Can humans understand the method?

Can the result generate new conjectures?

Can the discovery accelerate neighboring fields?

These questions concern mathematics itself.

They should not disappear beneath arguments over whether the discovery enhances the prestige of OpenAI, another AI company, a university, an individual researcher, or an established mathematical school.

OpenAI's August 2026 report illustrates the scale of the emerging opportunity. The company described ten results covering sphere packing, coding theory, non-sofic groups, operator algebras, arithmetic circuit complexity, quantum complexity, lattice cryptography, convex geometry, Ramsey theory, and extremal graph theory. According to OpenAI, the arguments were generated using an internal model, prepared into manuscripts with human involvement, and then formalized into Lean certificates.

The reported Navier–Stokes effort went still further. OpenAI states that it used coordinated groups of agents with access to code and cached internet material, generated millions of agent messages, shifted computational resources between mathematical problems, cross-pollinated intermediate discoveries between agent groups, and used later model versions as training continued during the project.

Whether every announced result ultimately survives the strongest mathematical scrutiny is not the central question of this paper.

The central fact is that a new form of mathematical actor has entered the research environment.

The scholarly infrastructure surrounding that actor remains primitive.


The Current Debate Collapses Different Questions

A significant weakness in contemporary arguments about AI and mathematical credit is the tendency to treat several independent propositions as though they were interchangeable.

They are not.

A user may have interacted with a model.

The user's interaction may have been eligible for model improvement.

Eligible data may have been included in some development process.

A later model may have descended from that development process.

The model may produce mathematics resembling something in the interaction.

None of those facts alone establishes that the interaction caused the mathematical discovery.

Likewise, absence of direct human access to a conversation does not by itself establish absence of indirect model exposure.

The relevant distinctions are permission, inclusion, exposure, derivation, causation, and attribution.

They must remain separate.

Consider the difference between these statements:

“OpenAI was contractually permitted to use this interaction for model improvement.”

“This interaction was actually included in training or evaluation data.”

“A model used in the later mathematical project was descended from a checkpoint influenced by that dataset.”

“The specific mathematical information within the interaction materially affected the model.”

“The later proof was intellectually derived from that information.”

“The original researcher therefore deserves scholarly attribution for the proof.”

Those are six different propositions.

A credible research-integrity framework cannot jump from the first to the sixth or from the second to the sixth without evidence connecting the intermediate stages.

This principle applies equally in the opposite direction.

A company cannot establish independent discovery merely by saying that its researchers never personally opened a particular user's conversation.

The relevant question is not merely whether a human employee viewed a source.

The relevant question is whether a demonstrable path of intellectual influence exists between the source and the resulting discovery.

That is a provenance question.


Terms of Service, Consent, and User Responsibility

One portion of the current controversy requires unusually careful treatment because legal permission and scholarly attribution are easily confused.

OpenAI's current Terms of Use for individual users outside Europe state that by using the services the user agrees to the terms. They further state that users retain ownership rights in their input, own output to the extent permitted by applicable law, and that OpenAI may use content to provide, maintain, develop, and improve its services. The terms also provide an opt-out mechanism for users who do not want their content used to train models.

OpenAI's current data-control documentation is even more explicit. For personal ChatGPT workspaces, conversations may be used to improve models unless the user disables the relevant training control; after opting out, new conversations are not used for model training. Business, Enterprise, Edu, and API offerings operate under different defaults, under which customer inputs and outputs are not used for model training unless the customer explicitly agrees.

This matters.

A user who voluntarily uses a service under terms permitting specified uses of submitted content cannot establish an ethical violation merely by later stating that the user did not personally understand or remember the provision.

Failure to read an available agreement does not, by itself, demonstrate that no agreement existed.

The familiar maxim that ignorance of the law is not an excuse is not a perfect legal analogy because statutes and private contracts are different legal instruments, and questions of contractual enforceability may depend on notice, assent, jurisdiction, consumer-protection law, and other circumstances. The stronger and more defensible principle is narrower:

A claim about unauthorized use must account for the agreement and data-control state that actually governed the interaction at the relevant time.

A researcher cannot simply assume privacy conditions inconsistent with the service configuration the researcher chose to use.

Researchers handling valuable unpublished mathematics therefore have responsibilities of their own.

If a person regards an unpublished proof, conjecture, construction, dataset, or technique as confidential intellectual material, that person should understand the data conditions of the computational environment into which the material is being entered.

This is ordinary information hygiene.

Researchers already distinguish between a public preprint server, private correspondence, an institutional repository, a confidential referee report, and a personal notebook. AI systems should be treated with the same seriousness.

But contractual permission is not the end of the argument.

Consent to model improvement does not automatically eliminate scholarly norms concerning acknowledgment.

Nor does it prove causal contribution.

A person may have authorized a mathematical conversation to participate in model improvement without thereby becoming an author of every later theorem produced by every descendant model.

Conversely, contractual permission to process material would not necessarily eliminate every ethical question if a later result could be shown to derive directly and recognizably from unpublished intellectual work.

Legal permission, causal provenance, and scholarly attribution remain different categories.

The virtue of a provenance system is that it prevents all three from being confused.


De-Identification Is Not Intellectual Provenance

The controversy also exposes an important technical misconception.

De-identification and provenance perform different functions.

De-identification attempts to reduce the connection between information and a particular identifiable person.

Provenance attempts to preserve the history and relationships through which information came to exist.

Removing a person's name from a mathematical interaction does not remove the mathematical structure contained within it.

A proof technique remains a proof technique.

A conjecture remains a conjecture.

A counterexample remains a counterexample.

A transformation remains a transformation.

Thus, the statement that material has been de-identified does not answer the intellectual question of where an idea came from.

But neither does the continued existence of the idea prove that a later model relied upon it.

This distinction is precisely why provenance must operate independently of personally identifying information.

A robust provenance architecture should be capable of stating that a cryptographically committed research object entered a particular dataset or process without necessarily exposing the identity of the contributor or the content of the research object itself.

Privacy and provenance are therefore not opposites.

Properly designed, provenance can strengthen privacy by allowing factual questions about lineage to be answered without unnecessary disclosure of the underlying material.


From File History to Intellectual Chain of Custody

The Personal Provenance Ledger proposed through Secretary Suite begins from a broader observation: modern information systems preserve enormous quantities of content while frequently losing the history that gives that content meaning.

A document may survive while its origin, permissions, source materials, revisions, rejected alternatives, transformations, contributors, and relationships to later documents become fragmented or disappear.

AI intensifies this weakness because information can move rapidly among people, models, applications, files, websites, tools, databases, and automated agents.

The solution is not simply better version history.

Version history tells us that an object changed.

Provenance tells us why it changed, what influenced the change, who or what performed it, what source objects participated, what permissions governed those sources, and what later objects descended from the transformation.

This concept is consistent with established provenance standards.

The World Wide Web Consortium's PROV architecture models provenance around entities, activities, and agents and explicitly represents relationships such as generation, usage, derivation, attribution, and influence. Its purpose is to allow provenance information generated by different systems to be represented and interchanged.

C2PA similarly demonstrates that cryptographically verifiable provenance can accompany digital content across transformations. Its Content Credentials architecture preserves assertions about an asset's origin and changes and uses digital signatures and content bindings to make provenance tamper-evident.

Crossref's relationship metadata approaches the problem from scholarly infrastructure, linking research objects so that publications, components, citations, versions, and related works can form what Crossref describes as a research nexus.

The required mathematical system should combine these traditions while extending them substantially.

Mathematical provenance must follow not merely files.

It must follow ideas through computational processes.


The Discovery Provenance Ledger

This paper proposes a specialized extension called the Discovery Provenance Ledger, or DPL.

The DPL would not attempt to store every private mathematical conversation in publicly readable form.

It would instead create a persistent chain of verifiable provenance assertions surrounding significant research events.

The fundamental object would be a provenance event.

A provenance event would identify an originating object or state, an activity performed upon it, the human or software agents involved, the resulting object or state, the time of the activity, the permissions governing the source, the system configuration under which the activity occurred, and cryptographic commitments sufficient to establish later that the recorded object has not been silently substituted.

A mathematical research interaction might therefore produce a record conceptually represented as

[ P_i = (E_{in}, A, G, T, M, R, C, E_{out}) ]

where E_{in} represents contributing entities, A the activity, G the responsible human or computational agents, T temporal information, M the model or computational environment, R the rights and permission state, C cryptographic commitments, and E_{out} the resulting entities.

Discovery provenance then becomes a graph:

[ G_P = (V,E) ]

in which vertices represent research objects, conversations, datasets, model checkpoints, computations, proofs, manuscripts, and other entities, while edges represent relationships such as used, generated, derived from, trained on, retrieved from, verified by, revised from, informed by, and published as.

The objective is not to declare ownership automatically.

The objective is to preserve evidence.


Five Levels of Provenance Evidence

The architecture should distinguish five increasingly strong forms of evidence.

The first is existence provenance. It establishes that a particular mathematical object or interaction existed at a particular time.

The second is eligibility provenance. It establishes whether that object was legally, contractually, or administratively eligible to participate in training, evaluation, retrieval, or another development process.

The third is exposure provenance. It establishes that a particular dataset, retrieval process, model-development activity, or computational system actually had access to the object or to a representation derived from it.

The fourth is derivation provenance. It establishes an identifiable transformation path connecting one research object to another.

The fifth is attribution provenance. It combines technical evidence with scholarly judgment to determine whether the demonstrated contribution is sufficiently substantive to warrant citation, acknowledgment, coauthorship, priority recognition, or another form of credit.

These levels prevent one of the most serious analytical errors in current AI debates:

Possible exposure is not demonstrated derivation.

And demonstrated derivation is not automatically authorship.


Training Provenance Is Especially Difficult

Neural models complicate provenance because learning does not resemble copying documents into a conventional database.

Model parameters are updated through large-scale optimization over enormous collections of examples.

Consequently, knowing that a training object was included in a dataset does not necessarily establish that a particular later model behavior can be causally attributed to that object.

This is a crucial limitation.

A provenance architecture must not promise more than the underlying science can establish.

A mathematically sophisticated framework should therefore distinguish dataset lineage from behavioral causation.

Dataset lineage can often be recorded deterministically.

If dataset D_1 contains object x, checkpoint M_1 was trained using D_1, and checkpoint M_2 descends from M_1, a provenance system can record those facts.

It cannot automatically conclude:

[ x \Rightarrow y ]

where y is some later theorem generated by M_2.

At most it has established a possible path of exposure.

The causal influence of individual training examples on a mature neural system may be diffuse, nonlinear, redundant, and difficult to isolate.

This uncertainty should itself become part of provenance.

Rather than pretending that causation is binary, a DPL could classify claims according to evidentiary confidence.

The system would thereby allow researchers to say:

“The model was demonstrably exposed to this material.”

or:

“The relevant checkpoint was demonstrably not trained on this material.”

without falsely asserting that exposure necessarily produced the final theorem.

That alone would resolve a substantial portion of contemporary disputes.


Training-State Attestation

A particularly valuable mechanism would be training-state attestation.

When a user submits content, the system should preserve a cryptographically verifiable record of the policy state governing that submission.

The record could include the service class, account category, training opt-in or opt-out state, temporary-chat state where applicable, relevant data-policy version, timestamp, and a cryptographic hash of the interaction or research object.

The user need not receive the provider's internal dataset architecture.

The provider need not publish the user's private content.

Both parties would nevertheless possess evidence of the governing state.

Years later, a disagreement would no longer depend upon statements such as:

“I thought training was disabled.”

“I assumed this was private.”

“We do not believe that account was eligible.”

“The setting may have changed.”

The relevant state would have been recorded when the event occurred.

This is especially important because OpenAI presently applies different training defaults to individual services and business-oriented services. A provenance system that ignores service class would therefore be incomplete by design.

Terms themselves should also be versioned.

A hash or persistent identifier for the exact governing policy should accompany the event.

Consent should have provenance too.


Model Lineage Must Be Recorded

Training-data provenance alone is insufficient.

Models themselves need lineage.

A mathematical discovery record should identify the exact model family, checkpoint, fine-tuning state, tool configuration, retrieval environment, system instructions where disclosure is permissible, and relevant computational orchestration.

This requirement becomes obvious when examining OpenAI's account of its September 2026 Navier–Stokes work.

OpenAI states that training of a new internal model began August 28, that the model was updated during the mathematical effort as further-trained versions became available, that agents were moved among problems, that intermediate Euler results were inserted into later prompts, and that useful insights were cross-pollinated among agent groups through Codex.

This is not a single inference.

It is an evolving computational research laboratory.

Traditional citation practices cannot adequately represent it.

The discovery itself has a lineage:

[ M_0 \rightarrow M_1 \rightarrow A_1 \rightarrow R_1 \rightarrow A_2 \rightarrow R_2 \rightarrow M_2 \rightarrow A_3 \rightarrow P ]

where models, agent populations, intermediate results, consolidation activities, and final proofs become successive provenance nodes.

If mathematical discovery increasingly occurs this way, preserving such chains will be essential to reproducibility.


Prompt Provenance

Prompts should also be recognized as research objects.

This does not mean that every sentence typed into an AI system deserves scholarly credit.

Most do not.

But prompts can contain substantive intellectual contributions.

A prompt may specify an unexplored reduction.

It may formulate a conjecture.

It may introduce a coordinate transformation.

It may instruct an AI system to combine previously unrelated mathematical frameworks.

It may identify a failed assumption.

It may provide a novel lemma.

It may direct thousands of agents toward a research strategy.

When prompts materially shape discovery, prompt provenance becomes intellectually relevant.

The proper question is not:

“Was there a prompt?”

There is always some computational instruction.

The proper question is:

What information did the prompt contribute that was necessary or materially useful to the resulting discovery?

The DPL should therefore preserve prompt hashes, timestamps, contributors, and derivational relationships while allowing sensitive prompt contents to remain sealed where appropriate.

A public paper might reveal only that a verified prompt existed and influenced a particular computational branch.

A trusted auditor could later inspect the sealed material during a dispute.

This creates accountability without requiring laboratories to expose every proprietary technique publicly.


Retrieval Provenance

Modern AI systems increasingly use search, browsing, databases, uploaded documents, vector stores, connected applications, cached websites, and specialized mathematical corpora.

Retrieval must therefore be recorded separately from training.

This distinction is indispensable.

If an AI agent retrieves a mathematician's preprint while solving a problem, that is a direct evidentiary relationship.

If the same preprint had been present somewhere in a trillion-token training corpus years earlier, the evidentiary relationship is radically different.

The first may be traceable at inference time.

The second may represent diffuse model exposure.

A provenance ledger should never collapse them.

Every substantive retrieval used during a research run should therefore generate a source relationship linking the retrieved object to the downstream reasoning activity.

Where a DOI, arXiv identifier, URL, dataset identifier, ORCID, or other persistent identifier exists, it should be captured.

Where no identifier exists, the system can generate a content hash and timestamp.

This would make AI-assisted scholarship more reproducible than much current human scholarship.

Instead of a model merely claiming that it “used sources,” the research record could show precisely which source objects entered which stages of the argument.


Intermediate Mathematical Objects Matter

One of the greatest potential contributions of AI provenance is preservation of intermediate mathematics.

Traditional papers frequently show polished results while omitting enormous amounts of intellectual exploration.

AI systems may generate orders of magnitude more intermediate reasoning.

Most will be useless.

Some will contain genuinely important mathematics that does not appear in the final proof.

A failed branch might contain a lemma useful elsewhere.

A rejected approach might expose a hidden equivalence.

A computational experiment may reveal a pattern leading to a future conjecture.

The provenance graph should therefore permit mathematically meaningful intermediate objects to survive independently of the final publication.

This creates something larger than a conventional paper.

It creates a discovery graph.

Research becomes navigable not only by publication but by intellectual ancestry.


Formal Verification as a Provenance Event

Formal proof systems such as Lean add another important layer.

Formalization should not be treated merely as an attachment demonstrating correctness.

It should be represented as a provenance transformation.

An informal mathematical argument becomes a formal object.

That transformation may expose hidden assumptions, repair gaps, clarify definitions, or reveal that the original argument cannot be formalized as stated.

The person or AI system responsible for formalization therefore participates in the discovery chain.

OpenAI's 2026 mathematical work is notable because the company reports using Lean certificates for the announced results, including formalization after the initial mathematical arguments were generated.

The chain should therefore distinguish:

informal discovery → manuscript preparation → formalization → machine verification → human interpretation → publication.

Each stage contributes different evidence.

Each deserves separate provenance.


Attribution Should Be Evidence-Based, Not Emotion-Based

This architecture leads to a difficult but necessary principle:

Resemblance is evidence worth investigating, not proof of appropriation.

Two mathematicians can independently discover the same theorem.

Two models can independently produce similar techniques.

A human and an AI can independently converge on the same argument because the mathematical structure strongly constrains possible solutions.

Priority claims therefore require more than conceptual similarity.

Likewise, a researcher who has repeatedly discussed a subject with an AI system cannot reasonably assume that every later system result in that subject descends from those discussions.

A user's expertise may make overlap more likely without establishing causal lineage.

The reverse error is equally serious.

An AI laboratory cannot dismiss legitimate concerns merely because neural-model influence is difficult to reconstruct.

Opacity should not become an evidentiary shield.

If laboratories create systems capable of consequential scientific discovery, they acquire a corresponding responsibility to preserve records sufficient to investigate discovery provenance.

NIST has already moved in this direction at a broader level. Its Generative AI Profile recommends documenting training-data sources so that the origin and provenance of generated content can be traced and emphasizes provenance as part of information integrity and accountability. NIST also notes that maintaining training-data provenance and supporting attribution of AI decisions to subsets of training data can assist transparency and accountability.

Scientific AI should demand at least that level of rigor.


Provenance Does Not Mean Surveillance

A predictable objection is that comprehensive provenance would require recording every researcher action permanently.

It should not.

A well-designed provenance architecture should follow principles of data minimization.

Not every private thought belongs in a public ledger.

Not every prompt belongs in a paper.

Not every model-development detail should become a trade-secret disclosure.

Not every researcher identity should be exposed.

Cryptographic systems allow evidence of existence and integrity to be separated from public disclosure.

A research object can be hashed.

A timestamp can be independently witnessed.

A dataset manifest can contain identifiers rather than raw material.

A private interaction can generate a sealed provenance receipt.

An auditor can verify an inclusion or exclusion claim without publishing the underlying content.

C2PA demonstrates the broader feasibility of attaching cryptographically verifiable provenance assertions to digital assets while allowing provenance systems to respect privacy and creator control.

The mathematical version should apply the same principle to intellectual lineage.


Privacy-Preserving Dispute Resolution

The strongest implementation would support privacy-preserving provenance queries.

Imagine that Researcher A claims that a private mathematical interaction influenced Model M.

The provider should not have to publish its complete training corpus.

Researcher A should not have to publish the confidential mathematics.

Instead, both sides could submit cryptographic commitments to a trusted provenance process.

The process could determine whether the committed research object, a sufficiently close derivative, or a designated interaction identifier appears in relevant dataset manifests.

The result might establish one of several facts:

No evidence of inclusion exists.

The interaction was eligible but not included.

The interaction was included in a development corpus.

A derivative representation was included.

The model checkpoint used in the disputed discovery predates the interaction.

The checkpoint descends from training that included the interaction.

The interaction entered a retrieval system during the research run.

A direct derivational relationship exists between an intermediate computational object and the disputed research object.

These are meaningful findings.

None requires an immediate conclusion regarding authorship.

They provide the evidence from which proper scholarly judgment can begin.


Provenance Receipts for Researchers

Researchers using AI systems should receive provenance receipts for significant work.

A receipt could state that on a particular date a hashed research object was submitted under a particular account class, that the training-control state was enabled or disabled, that the applicable policy version had a specified identifier, that a specified model processed the object, and that resulting output objects were generated.

Researchers could preserve these receipts in institutional repositories.

A future manuscript could reference them.

A patent dispute could reference them.

A priority dispute could reference them.

A journal editor could inspect them.

An AI provider could use corresponding records to demonstrate independence.

This would transform provenance from an abstract policy promise into a practical scholarly instrument.


Providers Benefit From Provenance Too

It would be a mistake to treat provenance as a burden imposed solely on AI companies.

AI laboratories have enormous incentives to adopt it.

Consider the current controversy.

OpenAI states that neither its researchers nor its agents accessed the unpublished Buckmaster–Alpöge work while developing its Navier–Stokes result, though the company says it cannot entirely rule out indirect contribution from de-identified usage data used to improve models.

Without stronger provenance, outsiders must decide whether they believe that explanation.

That is an unnecessarily weak position for everyone.

With strong provenance, OpenAI could potentially demonstrate that a disputed interaction was not present in the relevant training lineage, that it occurred after a checkpoint was trained, that it was excluded through the user's data-control state, or that the relevant agent environment never retrieved the disputed material.

Provenance therefore protects legitimate independent discovery.

It protects AI systems from false accusations just as surely as it protects human researchers from uncredited appropriation.

Good provenance is not anti-AI.

It is one of the technologies that will allow powerful AI research to remain credible.


Researchers Benefit From Provenance

The same system protects mathematicians.

A researcher who develops a novel idea privately can create a timestamped cryptographic commitment before publication.

The commitment proves existence without revealing the idea.

If the researcher later submits the idea to an AI system, a provenance event records that interaction.

If the idea later appears elsewhere, a clear chronological record exists.

The researcher no longer has to rely on screenshots, memory, email fragments, or public accusations.

The evidence speaks first.

This may be particularly important as AI systems become capable of independently exploring more mathematical territory.

Ironically, provenance may become more important precisely because genuine independent rediscovery becomes more common.

Without provenance, independent convergence and intellectual borrowing can look nearly identical.


Journals and Publishers Will Need New Standards

Scientific publishing should eventually require AI provenance statements for computationally significant discoveries.

A simple declaration that “AI was used” is too crude.

There is an enormous difference between using an AI system to correct grammar and using 10,000 reasoning agents to generate the central proof.

Disclosure should therefore describe functional contribution.

A journal should be able to determine whether AI systems performed literature retrieval, conjecture generation, proof search, computational experimentation, formal verification, manuscript drafting, or independent discovery.

Persistent relationships between research objects should then be deposited alongside conventional bibliographic metadata.

Crossref's existing relationship metadata and broader research-nexus concept provide useful infrastructure on which such relationships could eventually be built.

The scientific record would then preserve not merely the final paper but the discoverable lineage surrounding it.


A New Vocabulary of Mathematical Contribution

AI-assisted mathematics will also require a richer vocabulary than “author” and “tool.”

An AI system that fixes punctuation is clearly a tool.

An AI system that independently generates the mathematical argument is performing something substantially different.

A human who merely initiates a run is also performing a different role from a human who develops the central lemma, selects the successful branch, interprets the result, repairs the proof, and formalizes it.

OpenAI itself acknowledged this tension in its August 2026 mathematical announcement, stating that attribution should honestly reflect how a result was produced and that representing an AI-generated proof as entirely human-authored would misrepresent the discovery process.

Provenance allows contribution roles to become factual rather than ceremonial.

The scholarly community can then decide how those roles should map to authorship conventions.

The ledger records what happened.

The community decides what credit those events deserve.

That separation is essential.


Discovery Belongs to a Larger Intellectual Ecosystem

No mathematical AI system begins from nothing.

It inherits centuries of human mathematics.

Its notation is human.

Its problem statements are human.

Its formal languages are human.

Its training literature is largely human.

Its evaluation standards are human.

Its mathematical culture is human.

Recognition of AI discovery therefore should not erase the civilization-scale inheritance on which the system depends.

But neither should that inheritance be used to deny genuine machine discovery.

Human mathematicians also inherit almost everything.

Originality has never required intellectual creation ex nihilo.

The relevant question has always been whether a new intellectual object emerged through a sufficiently distinct act of reasoning or synthesis.

AI simply forces us to examine that question more carefully.

Provenance provides the evidence necessary to do so.


From Competition to Cooperative Mathematics

The most unfortunate possible outcome of current disputes would be a more secretive mathematical culture.

If mathematicians become afraid to discuss incomplete ideas with AI systems, colleagues, or public communities because every interaction is perceived as a potential race for ownership, mathematics loses.

If AI companies become afraid to investigate open problems because any overlap with human work invites accusations, mathematics loses again.

A provenance infrastructure offers a better outcome.

Humans and AI systems can collaborate aggressively while preserving intellectual lineage.

Researchers can share ideas with confidence because provenance follows them.

AI laboratories can explore mathematics confidently because independent discovery can be documented.

Journals can evaluate contributions with better evidence.

Future historians of mathematics can reconstruct how discoveries occurred.

The result is not less openness.

It is safer openness.


The Principle of Evidentiary Symmetry

A healthy system should impose the same evidentiary standard on every participant.

Researchers should not accuse laboratories of intellectual appropriation solely because a later result resembles their own unpublished thinking.

Laboratories should not dismiss researchers' concerns solely because the internal mechanics of training are inaccessible to outsiders.

Both claims should be testable against provenance.

This is the Principle of Evidentiary Symmetry:

[ \text{Strength of Attribution Claim} \propto \text{Strength of Provenance Evidence} ]

A strong accusation requires strong provenance.

A strong denial should likewise be supported by strong provenance when the relevant records are uniquely under the denying party's control.

The burden need not be absolute.

It needs to be evidentiary.


The Provenance of Consent

Consent itself deserves special emphasis.

Digital systems commonly treat consent as a momentary checkbox.

That is inadequate for long-lived intellectual workflows.

A provenance-aware system should treat consent as a versioned state.

If a researcher permits model training in January, disables it in March, uses a temporary research environment in April, and enters a business workspace in June, the provenance graph should preserve those distinctions.

The question “Did the researcher consent?” may therefore be badly formed.

The correct question is:

What permission state governed this specific object at this specific time in this specific service environment?

That is answerable.

And once it is answerable, much of the rhetorical confusion surrounding consent disappears.


A Standard for Claims of AI-Derived Discovery

Before a major AI laboratory announces a result as independently machine-generated, it should be capable of providing a provenance statement describing the relevant system and research chain.

Such a statement need not reveal proprietary model weights or private user data.

It should establish the problem source, model lineage, major retrieval sources, computational orchestration, human interventions, formal-verification process, and known limitations concerning training-data attribution.

Where complete tracing of training influence is scientifically impossible, the statement should say so.

Transparency does not require pretending uncertainty has vanished.

It requires precisely identifying where uncertainty remains.

That is the distinction between scientific reporting and public relations.


A Standard for Human Priority Claims

Human researchers should meet an analogous standard.

A credible priority or appropriation claim should identify a specific earlier intellectual object, establish its existence before the disputed discovery, identify a plausible exposure pathway, and demonstrate substantive mathematical correspondence beyond what would reasonably be expected through independent convergence or common knowledge.

Statements such as “the model understood our techniques unusually well” may justify investigation.

They do not, by themselves, establish intellectual descent.

Likewise, previous use of ChatGPT or another AI system does not automatically create an attribution interest in subsequent model behavior.

Such a rule would make AI-assisted research impossible.

Millions of users continuously contribute questions, corrections, preferences, examples, evaluations, and domain knowledge to interactive systems. Model improvement is necessarily collective at enormous scale.

If every interaction generated indefinite authorship rights over every downstream capability, no general-purpose learning system could function coherently.

The correct mechanism is provenance-based contribution analysis.


Provenance Before Accusation

This yields a simple scholarly norm:

Investigate lineage before alleging appropriation.

The standard protects both sides.

It discourages performative controversy.

It encourages technical evidence.

It shifts discussion from institutional reputation to reproducible facts.

It allows legitimate wrongdoing, if it occurs, to be demonstrated more convincingly.

And it allows legitimate independent discovery to be recognized more confidently.

Mathematics benefits either way.


Provenance as Scientific Infrastructure

The larger importance of this proposal extends well beyond present disputes.

AI-assisted science will increasingly involve molecular design, physics, materials science, biology, engineering, medicine, economics, climate modeling, and other fields in which ideas pass through increasingly complicated human–machine systems.

The provenance problem will follow them all.

NIST has already identified provenance tracking as important to generative-AI information integrity and risk management. W3C PROV supplies a mature ontology for describing entities, activities, and agents. C2PA demonstrates cryptographically protected content provenance. Crossref provides persistent scholarly relationships. Secretary Suite's Personal Provenance Ledger proposes extending chain of custody across ordinary digital actions.

What remains is to connect these ideas into a system explicitly designed for intellectual discovery.

Mathematics is an ideal place to begin.

Mathematical objects are unusually precise.

Formal verification provides unusually strong validation.

Research communities already care deeply about priority and attribution.

The output of a provenance experiment can therefore be evaluated with unusual rigor.


Conclusion

Artificial intelligence is becoming a participant in mathematical discovery.

That development should be approached with excitement rather than fear.

The possibility that computational systems can discover structures humans have not yet found is not an insult to mathematics.

It is an expansion of mathematics.

But extraordinary new discovery mechanisms require equally sophisticated mechanisms of intellectual accountability.

The present system is inadequate.

Screenshots are not provenance.

Memory is not provenance.

Similarity is not provenance.

De-identification is not provenance.

Terms of service are not provenance.

A public denial is not provenance.

An accusation is not provenance.

Provenance is the demonstrable chain connecting intellectual objects, permissions, people, systems, models, sources, transformations, and results.

Users have responsibilities within that chain. Researchers working with valuable unpublished material should understand the conditions of the systems they use. Where service terms expressly permit content to contribute to model improvement and an opt-out mechanism exists, failure to understand those conditions cannot by itself establish that the subsequent use was unauthorized. At the same time, contractual permission does not establish that a particular interaction actually entered training, does not establish causal influence, and does not resolve scholarly attribution.

AI providers have corresponding responsibilities. Systems capable of major scientific discovery should preserve enough internal lineage to investigate credible claims of intellectual descent without demanding blind trust from the scientific community.

The objective should not be to construct an ownership bureaucracy around every mathematical thought.

The objective is the opposite.

It is to make collaboration easier because intellectual history no longer disappears.

A researcher should be able to contribute without fearing that provenance will vanish.

An AI laboratory should be able to demonstrate independent discovery without asking the world to accept unsupported assurances.

A journal should be able to evaluate contribution without guessing.

A future mathematician should be able to inspect not merely the final theorem but the path through which the theorem entered human knowledge.

This leads to the governing principle of the proposed architecture:

Celebrate the discovery. Preserve the provenance. Attribute according to the evidence.

Mathematics does not benefit when human beings and machines fight over shadows.

It benefits when the provenance is clear enough that everyone can return to the mathematics.

And ultimately that should be the purpose of the entire enterprise:

not ownership of discovery,

but more discovery.


References

  1. OpenAI. “Ten Advances in Mathematics and Theoretical Computer Science.” August 1, 2026. OpenAI reported ten model-generated results spanning multiple mathematical fields and described subsequent human manuscript preparation and Lean formalization.

  2. OpenAI. “On the Navier–Stokes Millennium Prize Problem.” September 8, 2026. Describes the multi-agent research process, model development, Lean verification, concurrent human work, and OpenAI's statement regarding specific and de-identified user data.

  3. Hart, Robert. “Mathematicians Want Proof OpenAI Didn't Use Their Work.” The Verge, September 10, 2026. Reports Andreas Thom's concerns regarding training-data provenance and OpenAI mathematical research.

  4. OpenAI. Terms of Use. Effective January 1, 2026. Defines user content, ownership, service-improvement rights, and training opt-out provisions applicable to individual users outside the EEA, Switzerland, and United Kingdom.

  5. OpenAI. “How Your Data Is Used to Improve Model Performance.” Updated March 13, 2026. Describes training practices for individual-user services and available opt-out mechanisms.

  6. OpenAI Help Center. Data Controls FAQ and related training-control documentation. Describes the “Improve the model for everyone” control and differences between individual and business-oriented services.

  7. OpenAI. Business Data Privacy, Security, and Compliance. States that organization data from Business, Enterprise, Edu, Healthcare, Teachers, and API offerings is not used for model training by default.

  8. Swygert, John. “The Personal Provenance Ledger: Automatic Chain of Custody for Every Meaningful Digital Action; A Secretary Suite Project.” Secretary Suite, August 2, 2026. Proposed architecture for persistent provenance across human and AI-assisted digital activity.

  9. World Wide Web Consortium. PROV-O: The PROV Ontology. W3C Recommendation, April 30, 2013. Establishes interoperable provenance concepts based upon entities, activities, agents, derivation, attribution, and influence.

  10. Groth, Paul, and Luc Moreau, eds. PROV-Overview: An Overview of the PROV Family of Documents. World Wide Web Consortium, April 30, 2013. Defines provenance as information regarding the entities, activities, and people involved in producing a data object or thing.

  11. Coalition for Content Provenance and Authenticity. C2PA Content Credentials Technical Specification, Version 2.4. Describes cryptographically verifiable and tamper-evident provenance assertions for digital content.

  12. Crossref. Relationship Metadata. Describes links among scholarly research objects and their role in constructing the research nexus.

  13. Autio, Chloe, Reva Schwartz, Jesse Dunietz, Shomik Jain, Martin Stanley, Elham Tabassi, Patrick Hall, and Kamie Roberts. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile. NIST AI 600-1, National Institute of Standards and Technology, 2024; updated 2026. Recommends training-data provenance documentation and provenance mechanisms for generative-AI information integrity. DOI: 10.6028/NIST.AI.600-1.


Copyright © John Swygert 2026
TSTOEAO.com
IvoryTowerJournal.com
SecretarySuite.com
Ivory Tower Publishing


No comments:

Post a Comment